BioPharmaTrend
Latest Insights
Companies
  • Companies Directory
  • Case Studies
Newsletter
About
  • At a Glance
  • Our Team
  • Advisory Board
  • Citations and Press Coverage
  • Partner Events Calendar
  • Advertise with Us
 
 Subscribe 
Sign in
  • AI in Bio
  • Tech Giants
  • Next-Gen Tools
  • Business Intelligence

  News

Basecamp Launches Trillion Gene Atlas Targeting Over 100 Million Species for Better AI Training Data

by Anastasiia Rohozianska   •   March 18, 2026

Disclaimer: All opinions expressed by Contributors are their own and do not represent those of their employers, or BiopharmaTrend.com.
Contributors are fully responsible for assuring they own any required copyright for any content they submit to BiopharmaTrend.com. This website and its owners shall not be liable for neither information and content submitted for publication by Contributors, nor its accuracy.

# AI in Bio   
Share:   Share in LinkedIn  Share in Bluesky  Share in Reddit  Share in Hacker News  Share in X  Share in Facebook

Basecamp Research launched the Trillion Gene Atlas, a program designed to expand known genetic diversity by collecting genomic data from over 100 million species across thousands of global sites. The initiative was launched in collaboration with Anthropic, Ultima Genomics, and PacBio, with additional support from NVIDIA’s AI computing infrastructure, and is intended to address a constraint in biological AI, where most models rely on relatively limited public repositories.

The Trillion Gene Atlas is structured around three components: large-scale sequencing, global data acquisition networks, and high-performance compute. 

  • Ultima Genomics contributes wafer-based short-read sequencing systems designed for large-scale throughput, with its recently launched UG200 capable of processing more than 60,000 human genomes per year,
  • PacBio provides long-read sequencing to preserve genomic context and enable higher-resolution reconstruction of complex samples.
  • Compute infrastructure is provided through NVIDIA systems, with software such as Parabricks used to accelerate metagenomic assembly and CUDA-X Libraries to process unstructured data.
  • Anthropic is collaborating to integrate its Claude models with Basecamp’s EDEN system to build an end-to-end workflow that links clinical data interpretation directly to AI-driven therapeutic design.

Basecamp estimates that processing workloads at the scale of quadrillions of DNA base pairs, which previously could take over 20 years, may be reduced to under two years through parallelized pipelines and optimized algorithms.

Addressing Limited Dataset Diversity in AI Drug Discovery

According to Basecamp Research, most sequence-based foundation models still depend on similar public repositories, with around 80% trained on datasets containing fewer than 250 million sequences.

As model scale and compute capacity grow, dataset diversity has become a key factor in advancing AI-driven drug development and benchmarking. Over six years, Basecamp has established biodiversity partnerships across 31 countries, using distributed sampling and off-grid sequencing systems to collect genomic material from environments that are typically underrepresented in existing datasets. 

The framework incorporates local infrastructure development and aligns with emerging regulations around digital sequence information and benefit sharing.

Basecamp’s recently launched foundation models, EDEN, were trained on a proprietary BaseData dataset, reportedly exceeding public resources in scale, including around 10 billion previously uncharacterized genes from roughly one million species. 

According to the company, increasing dataset diversity enabled shifts in model behavior, moving from predictive tasks toward generating therapeutic candidates directly from disease-related inputs.

In internal validation experiments, EDEN reportedly demonstrated activity in primary human T cells without prior exposure to human or clinical datasets. The system has also been used to generate antimicrobial peptides, with reported hit rates of 97% against selected priority pathogens, and to explore gene insertion approaches under what the company terms AI-programmable gene insertion.


Basecamp’s Trillion Gene Atlas enters a broader push toward building integrated stacks that span data generation, model training, and experimental validation—an emerging infrastructure layer we examine in more detail in our deep dive on platforms for human-relevant drug discovery.

Resources such as UK Biobank, alongside multi-million-cell initiatives including the Human Cell Atlas and the Tahoe & Arc Institute’s Virtual Cell Atlas, are increasingly used as reference frameworks for mapping tissues, cell states, and disease biology, supporting target identification and model development across therapeutic areas.

Earlier this year, Illumina launched the Billion Cell Atlas with AstraZeneca, Merck, and Eli Lilly, a CRISPR perturbation dataset designed to train models on biological responses across more than 1 billion individual cells and 200 disease-relevant cell lines.

Just days before Basecamp’s announcement, Roche unveiled its own NVIDIA-based AI factory with more than 3,500 GPUs to support large-scale model training, biological data analysis, and simulation-heavy workflows across drug discovery and manufacturing. 

Further developments are expected from Eli Lilly, which has announced plans to build a $1 billion NVIDIA-backed AI research hub in San Francisco, alongside its recently launched LillyPod AI supercomputer to support drug development and manufacturing workflows.

In the UK, Arctoris recently launched their Biophysics Centre of Excellence, expanding automated lab infrastructure to generate high-resolution molecular interaction data at scale, targeting a growing bottleneck where AI-designed compounds outpace the ability of labs to test them.

Topic: AI in Bio

Share:   Share in LinkedIn  Share in Bluesky  Share in Reddit  Share in Hacker News  Share in X  Share in Facebook

You may also be interested to read:

Illumina Launches Billion-Cell CRISPR Dataset for AI-Driven Drug Discovery, Built with AstraZeneca, Merck, and Lilly
by Anastasiia Rohozianska
Anthropic Launches Claude for Healthcare at JPM26
by Anastasiia Rohozianska
Xaira Publishes Largest Public Perturb-seq Atlas to Advance Virtual Cell Modeling
by Roman Kasianov
Basecamp Research Partners with Cameroon to Share Benefits from Genetic Data Use in Landmark Agreement
by Roman Kasianov
Basecamp Research Introduces ZymCTRL: An Open-Source AI Tool for Enzyme Design
by Roman Kasianov
Basecamp Research Introduces BaseFold, Enhancing Accuracy in Protein Structure Prediction
by Andrii Buvailo, PhD

 

BiopharmaTrend.com

Where Tech Meets Bio
mail  Newsletter
in  LinkedIn
x  X
rss  RSS Feed

About


  • What we do
  • Press & Citations
  • Terms of Use
  • Privacy Policy
  • Cookies Policy
  • Disclaimer

Topics


  • News
  • AI in Bio
  • Tech Giants
  • Next-Gen Tools

Explore


  • Premium Insights
  • Business Intelligence
  • Companies
  • Events
  • Authors

Partner


  • Sponsorship
  • Editorial Calendar

© WTMB Research & Media, S.L. (WTMB Group)   2026
We use cookies to personalise content and to analyse our traffic. You consent to our cookies if you continue to use our website. Read more details in our cookies policy.