Basecamp Launches Trillion Gene Atlas Targeting Over 100 Million Species for Better AI Training Data
Basecamp Research launched the Trillion Gene Atlas, a program designed to expand known genetic diversity by collecting genomic data from over 100 million species across thousands of global sites. The initiative was launched in collaboration with Anthropic, Ultima Genomics, and PacBio, with additional support from NVIDIA’s AI computing infrastructure, and is intended to address a constraint in biological AI, where most models rely on relatively limited public repositories.
The Trillion Gene Atlas is structured around three components: large-scale sequencing, global data acquisition networks, and high-performance compute.
- Ultima Genomics contributes wafer-based short-read sequencing systems designed for large-scale throughput, with its recently launched UG200 capable of processing more than 60,000 human genomes per year,
- PacBio provides long-read sequencing to preserve genomic context and enable higher-resolution reconstruction of complex samples.
- Compute infrastructure is provided through NVIDIA systems, with software such as Parabricks used to accelerate metagenomic assembly and CUDA-X Libraries to process unstructured data.
- Anthropic is collaborating to integrate its Claude models with Basecamp’s EDEN system to build an end-to-end workflow that links clinical data interpretation directly to AI-driven therapeutic design.
Basecamp estimates that processing workloads at the scale of quadrillions of DNA base pairs, which previously could take over 20 years, may be reduced to under two years through parallelized pipelines and optimized algorithms.
Addressing Limited Dataset Diversity in AI Drug Discovery
According to Basecamp Research, most sequence-based foundation models still depend on similar public repositories, with around 80% trained on datasets containing fewer than 250 million sequences.
As model scale and compute capacity grow, dataset diversity has become a key factor in advancing AI-driven drug development and benchmarking. Over six years, Basecamp has established biodiversity partnerships across 31 countries, using distributed sampling and off-grid sequencing systems to collect genomic material from environments that are typically underrepresented in existing datasets.
The framework incorporates local infrastructure development and aligns with emerging regulations around digital sequence information and benefit sharing.
Basecamp’s recently launched foundation models, EDEN, were trained on a proprietary BaseData dataset, reportedly exceeding public resources in scale, including around 10 billion previously uncharacterized genes from roughly one million species.
According to the company, increasing dataset diversity enabled shifts in model behavior, moving from predictive tasks toward generating therapeutic candidates directly from disease-related inputs.
In internal validation experiments, EDEN reportedly demonstrated activity in primary human T cells without prior exposure to human or clinical datasets. The system has also been used to generate antimicrobial peptides, with reported hit rates of 97% against selected priority pathogens, and to explore gene insertion approaches under what the company terms AI-programmable gene insertion.
Basecamp’s Trillion Gene Atlas enters a broader push toward building integrated stacks that span data generation, model training, and experimental validation—an emerging infrastructure layer we examine in more detail in our deep dive on platforms for human-relevant drug discovery.
Resources such as UK Biobank, alongside multi-million-cell initiatives including the Human Cell Atlas and the Tahoe & Arc Institute’s Virtual Cell Atlas, are increasingly used as reference frameworks for mapping tissues, cell states, and disease biology, supporting target identification and model development across therapeutic areas.
Earlier this year, Illumina launched the Billion Cell Atlas with AstraZeneca, Merck, and Eli Lilly, a CRISPR perturbation dataset designed to train models on biological responses across more than 1 billion individual cells and 200 disease-relevant cell lines.
Just days before Basecamp’s announcement, Roche unveiled its own NVIDIA-based AI factory with more than 3,500 GPUs to support large-scale model training, biological data analysis, and simulation-heavy workflows across drug discovery and manufacturing.
Further developments are expected from Eli Lilly, which has announced plans to build a $1 billion NVIDIA-backed AI research hub in San Francisco, alongside its recently launched LillyPod AI supercomputer to support drug development and manufacturing workflows.
In the UK, Arctoris recently launched their Biophysics Centre of Excellence, expanding automated lab infrastructure to generate high-resolution molecular interaction data at scale, targeting a growing bottleneck where AI-designed compounds outpace the ability of labs to test them.
Topic: AI in Bio