BioPharmaTrend
Latest Insights
Companies
  • Companies Directory
  • Case Studies
Newsletter
About
  • At a Glance
  • Our Team
  • Advisory Board
  • Citations and Press Coverage
  • Partner Events Calendar
  • Advertise with Us
 
 Subscribe 
Sign in
  • AI in Bio
  • Tech Giants
  • Next-Gen Tools
  • Business Intelligence

  Business Intellilgence

Into the Dark: Finding Novel Drug Targets Within the Depths of Our Proteome

by Louise von Stechow  (contributor ) , Sudhakaran Prabakaran, PhD  (contributor )   •   Aug. 31, 2026

Disclaimer: All opinions expressed by Contributors are their own and do not represent those of their employers, or BiopharmaTrend.com.
Contributors are fully responsible for assuring they own any required copyright for any content they submit to BiopharmaTrend.com. This website and its owners shall not be liable for neither information and content submitted for publication by Contributors, nor its accuracy.

Share:   Share in LinkedIn  Share in Bluesky  Share in Reddit  Share in Hacker News  Share in X  Share in Facebook

Drug discovery has never had a richer therapeutic toolbox. Antibody-drug conjugates, antisense oligonucleotides, CRISPR therapeutics, RNA medicines, CAR-T cells, PROTACs, and molecular glues have transformed how medicine can manipulate biology. Yet every modality still depends on the same prerequisite: a biologically meaningful target. Many of the roughly 20,000 canonical human genes and their protein products are poor therapeutic candidates because they lack disease relevance, cannot be addressed by certain modalities, or cause unacceptable side effects when manipulated.

Perhaps the limiting factor is not only how we drug biology, but how much biology we know.

In May 2026, the TransCODE consortium put numbers to the problem. After screening 95,520 proteomics experiments against 7,264 non-canonical open reading frames, it found that roughly a quarter—about 1,785—gave rise to detectable protein-like molecules, which the consortium formally named “peptideins.” These entities make up the so-called “dark proteome.” While the consortium contributed scale and standardization, the underlying biology was not new. Over the past decade, work from various groups, including Prabakaran et al. (2014), Martinez et al. (2020), and Chen et al. (2020), has shown that these once-“non-coding” regions encode real, functional, and druggable proteins.

In a 2021 pan-cancer analysis, Erady et al. conducted an in-silico screen for small-molecule inhibitors targeting nORF-encoded proteins dysregulated in stomach and esophageal cancer. The authors showed that these targets could, in principle, be drugged. What matters for drug discovery is this: the option space is not only larger than we assumed; it is tractable. A number of biotechs are now exploring ways to target dark proteins systematically.

 

A genome that did not behave as advertised

If drug discovery is limited by the number of biological targets available, then understanding the true size of the functional proteome becomes one of the most important questions in biotechnology.

For decades, molecular biology taught a simple view of the human genome. About 2% of it was thought to encode proteins, while the remaining 98% was widely dismissed as “junk DNA.” The Human Genome Project, completed in 2003, did not settle that question; it sharpened it. Researchers found that the protein-coding fraction was even smaller than expected—closer to 1.5%—and that the remaining genome contained regulatory elements, transposable elements, pseudogenes, long non-coding RNAs, and other biologically active features.

More importantly, the term “non-coding” increasingly appeared to reflect the limits of technology and annotation rather than the limits of biology. Some regions long classified as non-coding turned out to produce not only regulatory RNAs but also previously overlooked proteins.

 

Where the dark proteome began

The realization that we had underestimated the size of the human proteome emerged gradually as new technologies exposed the limits of existing reference maps. Non-canonical proteins had remained hidden from classical proteomics methods because of their short half-lives, context-dependent expression, and low abundance. Even when detected, peptides that did not match reference databases were often discarded as noise.

A 2014 Nature Communications paper provided the first systematic evidence that RNAs labeled “non-coding” were producing measurable proteins. Using quantitative mass spectrometry, Prabakaran et al. showed that ribosomes were translating supposedly non-coding regions into reproducible, biologically meaningful peptides across multiple cell types. The study received modest attention at the time but, in retrospect, it began the process of recharting the boundaries of the human proteome.

Over the following decade, researchers built the evidence needed to convince the field. Ribosome profiling revealed active translation, targeted mass spectrometry confirmed the proteins, immunopeptidomics showed that dark peptides were presented on HLA molecules, and genetic studies linked non-canonical coding regions to cancer, rare diseases, infectious diseases, and neuropsychiatric disorders.

In 2026, TransCODE gave the field a shared vocabulary and classification system for microproteins and peptideins. The question is no longer whether the dark proteome exists. It is which of these proteins should be drugged and how.

 

From hidden protein to cancer target

One of the clearest examples is c10riboseqorf92, a 123-amino-acid microprotein encoded within the long non-coding RNA OLMALINC, also known as LINC00263. According to standard annotation, this region should not have produced a protein at all. Until recently, it was little more than an unexplained ribosome footprint in a public dataset. Today, it is among the most rigorously interrogated candidates to emerge from the dark proteome.

Its assessment illustrates how the field now evaluates newly discovered ORFs. Ribosome profiling identified active translation of the open reading frame, and targeted mass spectrometry confirmed that a product was made. CRISPR knockout data from the Dependency Map showed that disrupting it impaired viability in 415 of 485 cancer cell lines, and rescue experiments confirmed that the effect came from the encoded microprotein itself rather than from the RNA transcript that encodes it. Co-essentiality analyses tied it to mitosis and the DNA-damage response, while single-cell RNA sequencing localized its expression to actively dividing cells.

Multiple independent lines of evidence converge on a genuine role in cancer cell survival. The consortium itself stops short of calling c10riboseqorf92 a bona fide protein—the formal protein-coding evidence is not yet in hand—and that caution is instructive: the phenotype is compelling even where annotation remains unsettled.

Notably, c10riboseqorf92 is only one of the 7,264 ncORFs the consortium analyzed, of which roughly 1,785 met the criteria for peptidein status. So which of these could be classified as a potential drug target? The discovery problem is becoming a prioritization problem.

 

Addressing the prioritization challenge with AI

While discovering a dark protein is no longer the hard part, deciding which candidates justify a decade of drug development—and which therapeutic modality is best suited to target them—is. A promising target must show disease relevance, genetic support, appropriate tissue expression, structural tractability, and compatibility with a therapeutic modality.

This is where artificial intelligence could have a major impact. The traditional hand-curated funnel, in which a biologist evaluates one target at a time, cannot scale to a proteome that may be an order of magnitude larger than previously assumed. Yet many standard prediction tools—including AlphaFold confidence scores, druggability algorithms, and pocket-detection models—were trained on canonical proteins and may behave unpredictably when applied to the short, disordered, lineage-specific proteins of the dark proteome.

Several biotechs, including ProFound Therapeutics, DeepTarget Bio, and NonExomics, are building AI platforms to integrate evidence for the biological relevance of non-canonical proteins in disease contexts. Others, including Enara Bio Therapeutics, RyboDyn, and Alithea Bio, are applying similar logic to cancer immunotherapy, combining computational approaches with wet-lab validation to identify dark antigens uniquely expressed by tumor cells.

Whether all newly discovered proteins will prove therapeutically relevant remains an open question. Many will have limited biological importance, and others will be impossible to target. But biology has repeatedly shown that expanding our understanding of living systems expands our ability to treat disease.

 

The next frontier of drug discovery

The field is at an inflection point. Discovery methods are mature, terminology is standardized, and the first targets — c10riboseqorf92 foremost among them — have been characterized in depth in the peer-reviewed literature. Patents are in place, and capital is committed.

The next test is in vivo. Many mouse models do not express human-specific dark proteins, requiring humanized models, xenografts, or engineered animals expressing the human ORF. Peptide-level pharmacodynamics will require targeted mass spectrometry with isotopically labeled standards. Lineage-specific peptideins may also evade central tolerance, making them either powerful vaccine antigens or risky biological targets, depending on the modality.

The next few years of in vivo data will define the field. But the core argument is already clear: drug discovery is no longer limited by modality. Over the past fifteen years, biotech has built an extraordinary modality toolkit. What it lacks are targets.

The dark proteome may supply the next decade of them. The companies that can systematically pair these targets with the right modalities will shape biotech in the era beyond the patent cliff.


 

References 

  1. International Human Genome Sequencing Consortium. (2004). Finishing the euchromatic sequence of the human genome. Nature. DOI: 10.1038/nature03001.
  2. Ingolia, N. T., Ghaemmaghami, S., Newman, J. R. S., & Weissman, J. S. (2009). Genome-wide analysis in vivo of translation with nucleotide resolution using ribosome profiling. Science. DOI: 10.1126/science.1168978.
  3. Germain, P.-L., Ratti, E., & Boem, F. (2014). Junk or functional DNA? ENCODE and the function controversy. Biology & Philosophy. DOI: 10.1007/s10539-014-9441-3.
  4. Prabakaran, S., Hemberg, M., Chauhan, R., et al. (2014). Quantitative profiling of peptides from RNAs classified as noncoding. Nature Communications. DOI: 10.1038/ncomms6429.
  5. Chi, K. R. (2016). The dark side of the human genome. Nature. DOI: 10.1038/538275a.
  6. Chen, J., Brunner, A.-D., Cogan, J. Z., et al. (2020). Pervasive functional translation of noncanonical human open reading frames. Science. DOI: 10.1126/science.aay0262.
  7. Martinez, T. F., Chu, Q., Donaldson, C., et al. (2020). Accurate annotation of human protein-coding small open reading frames. Nature Chemical Biology. DOI: 10.1038/s41589-019-0425-0.
  8. Neville, M. D. C., Kohze, R., Erady, C., et al. (2021). A platform for curated products from novel open reading frames prompts reinterpretation of disease variants. Genome Research. DOI: 10.1101/gr.263202.120.
  9. Prensner, J. R., Enache, O. M., Luria, V., et al. (2021). Noncanonical open reading frames encode functional proteins essential for cancer cell survival. Nature Biotechnology. DOI: 10.1038/s41587-020-00806-2.
  10. Mudge, J. M., Ruiz-Orera, J., Prensner, J. R., et al. (2022). Standardized annotation of translated open reading frames. Nature Biotechnology. DOI: 10.1038/s41587-022-01369-0.
  11. Mohsen, J. J., Martel, A. A., & Slavoff, S. A. (2023). Microproteins—Discovery, structure, and function. Proteomics. DOI: 10.1002/pmic.202100211.
  12. Ge, A., Chan, C., & Yang, X. (2024). Exploring the dark matter of human proteome: The emerging role of non-canonical open reading frame (ncORF) in cancer diagnosis, biology, and therapy. Cancers. DOI: 10.3390/cancers16152660.
  13. Hofman, D. A., Ruiz-Orera, J., Yannuzzi, I., et al. (2024). Translation of non-canonical open reading frames as a cancer cell survival mechanism in childhood medulloblastoma. Molecular Cell. DOI: 10.1016/j.molcel.2023.12.003.
  14. Cardon, T., Fournier, I., & Salzet, M. (2025). Chasing the ghost proteome in the dark matter. Molecular & Cellular Proteomics. DOI: 10.1016/j.mcpro.2025.101076.
  15. Rajinikanth, N., Chauhan, R., & Prabakaran, S. (2025). Harnessing noncanonical proteins for next-generation drug discovery and diagnosis. WIREs Mechanisms of Disease, 17, e70001. DOI: 10.1002/wsbm.70001
  16. Deutsch, E. W., Kok, L. W., Mudge, J. M., et al. (2026). Expanding the human proteome with microproteins and peptideins. Nature. DOI: 10.1038/s41586-026-10459-x.

 

Disclosure:

Sudhakaran Prabakaran is co-founder and chief executive officer of NonExomics.

Share:   Share in LinkedIn  Share in Bluesky  Share in Reddit  Share in Hacker News  Share in X  Share in Facebook

BiopharmaTrend.com

Where Tech Meets Bio
mail  Newsletter
in  LinkedIn
x  X
rss  RSS Feed

About


  • What we do
  • Press & Citations
  • Terms of Use
  • Privacy Policy
  • Cookies Policy
  • Disclaimer

Topics


  • News
  • AI in Bio
  • Tech Giants
  • Next-Gen Tools

Explore


  • Premium Insights
  • Business Intelligence
  • Companies
  • Events
  • Authors

Partner


  • Sponsorship
  • Editorial Calendar

© WTMB Research & Media, S.L. (WTMB Group)   2026
We use cookies to personalise content and to analyse our traffic. You consent to our cookies if you continue to use our website. Read more details in our cookies policy.