Generare Raises €20M to Explore Uncharacterized Chemical Space from Microbial DNA
Paris-based techbio Generare has raised €20 million in a Series A round co-led by Alven and daphni, with participation from existing investors. The company is focused on generating previously inaccessible small-molecule data by decoding microbial genomes, aiming to expand the pool of compounds available for drug discovery.
The funding will support scaling its dataset, platform capacity, and team, with a target of increasing output to thousands of new molecules over the coming years.
Generare, founded in 2023 by Guillaume Vandenesch and Vincent Libis, is building a discovery platform centered on microbial DNA as a source of new chemistry. Recent reviews note that more than half of approved drugs are related to natural products, and microbial metabolites remain a major source of chemically novel scaffolds for therapeutics. At the same time, the majority of naturally encoded molecules remain uncharacterized.
Most screening libraries are derived from a narrow subset of known chemistry, often cited at around 3% of accessible chemical space. This existing subset has produced hundreds of approved drugs, yet it also defines the boundaries of what current machine learning systems can generalize over. As a result, many AI-driven pipelines iterate within the same regions of chemical space, even when architectures differ.
Generare’s platform combines high-throughput cloning, sequencing, and expression systems to systematically explore microbial genomes. It identifies biosynthetic gene clusters likely to encode bioactive compounds, expresses them in laboratory systems, and characterizes resulting molecules for structure and biological activity. Each cycle produces new small molecules and associated data, which are added to a growing proprietary dataset.

Image credit: Generare
Over time, this creates a feedback loop where better data leads to more precise discovery, and more precise discovery expands the dataset further. The result is a system focused on generating entirely new molecular starting points, derived from evolutionary processes and absent from public databases or existing AI training sets.
Generare reports generating over 200 previously uncharacterized small molecules per cycle, compared to an industry baseline of roughly 45. These outputs are positioned as inputs for downstream drug discovery programs, particularly for small-molecule therapeutics.
The Series A funding will be used to expand this dataset and increase throughput. Generare aims to scale its output to more than 2,000 molecules by 2027, with longer-term targets exceeding 10,000. The company currently employs around 25 staff across computational biology, chemistry, and synthetic biology, and plans to expand its team alongside platform scaling.
Other companies and research groups are also addressing the limitation of fresh, large-scale training data for AI:
- Illumina launched its Billion Cell Atlas, describing it as a genome-wide perturbation dataset built to support target validation and AI training.
- Arc Institute’s Virtual Cell Atlas now aggregates data from more than 600 million cells across large perturbation experiments.
- Bioptimus recently launched STELA, a multimodal tissue atlas spanning three continents and targeting up to 100,000 patient samples
- Basecamp Research launched the Trillion Gene Atlas, aiming to sequence genetic material from 100 million species.
All those efforts support the idea that progress in AI for drug discovery is increasingly tied to the scale, diversity, and structure of underlying datasets.
While Generare is focused on generating new molecular starting points, similar platforms are already pushing candidates into development. Enveda, for example, has advanced a nature-derived small molecule from its AI discovery system into the clinic, with positive clinical data from phase 1b.
Topic: Biotech Ventures