Insider Brief
- Ambitious Bio launched Corpus, a healthy-human-tissue protein map that the company says increased detected protein inventories by 28% to 78% across 181 matched comparisons with public references and published studies.
- In a comparison with the Human Protein Atlas, Corpus identified 56% more proteins in the median shared tissue and reported healthy-tissue expression for some drug targets not previously detected in that reference.
- The company says the data could help drug developers assess potential off-target safety risks earlier by showing where disease-related protein targets also appear in healthy tissues.
- Image: Ambitious turns matter into model-ready data, powering biological AI. (Ambitious Bio)
PRESS RELEASE — Ambitious Bio (“Ambitious”), a life sciences and AI infrastructure company creating a systematic record of human biology for AI, announced today the launch of Corpus™, the world’s most complete map of protein presence and relative abundance across healthy human tissues. In 181 head-to-head tissue comparisons spanning the major public references and published studies, adding Corpus increased the detected protein inventory every time, with the gain in a typical matched tissue ranging from 28% to 78%.
AI and biopharma leaders increasingly recognize both the opportunity and the constraint: Biology could become one of AI’s largest markets beyond coding, but much of the human data needed to realize that promise has never been systematically generated. Corpus demonstrates both the scale of that missing-data problem and Ambitious’s ability to build the systems that will eliminate it. This ability matters: The companies and countries that can generate reliable biological measurements at scale will be positioned to lead biological AI, and control the medicines and industries that emerge from it.
“We have complete bills of material for cars, wristwatches and toasters, down to the last screw. We still don’t have the same for the human body,” said Elizabeth Hudson, founder and CEO of Ambitious. “As AI becomes more capable of designing interventions into biology, the value of a more complete map of the system it’s trying to change only increases.”
Innovation Depends On Better Data
ChatGPT and Cowork showed what happens when a bottleneck between technical capability and widespread use is removed. Their impact was also predicated on a vast preexisting record of human language. Biology begins one step earlier: Much of human biology has never been systematically measured, and those data cannot simply be scraped from an existing record. They must be generated from physical specimens and made comparable across people, tissues, states of health and time. A ChatGPT or Cowork-scale unlock in biological AI therefore depends first on building the missing data foundation.
Backed by $6 million in seed funding, Ambitious was built for this problem. Rather than trying to extract value by happenstance from the data exhaust generated by existing clinical systems, the company works backward from a consequential decision to determine what evidence it requires. It then builds the infrastructure needed to generate that evidence repeatedly and at scale.
Starting With the Body’s Incomplete Protein Parts List
For its first product, Ambitious chose one high-value intersection: healthy human tissue measured at the protein layer. The genome is a parts catalog—it describes what the body could make—but proteins are the working parts it actually produces, in particular places and at particular levels. Many medicines work by engaging a particular protein, their target. The cleanest targets are abundant in diseased tissue but scarce or absent in vital healthy tissues. When that perfect separation does not exist, developers need to know how wide the margin is – where else the target appears and at what levels. Healthy tissue therefore provides the baseline scientists need to interpret what disease has changed and identify where a medicine might act unintentionally.
As one example of Corpus’s scale and quality, Ambitious released findings from a matched comparison with the Human Protein Atlas (HPA), the field’s most widely used reference. The comparison shows Corpus expanded HPA’s protein parts list by 56% in the median shared tissue (Figure 1); 56 proteins that hadn’t previously been identified as present in that tissue for every 100 already known. Ambitious repeated this analysis with the next six largest references, revealing 79% to 138% more protein types in the median tissue not previously visible to researchers. This added visibility carries significant insights for AI and drug researchers to reason on.
A systematic map of biology is a prerequisite for AGI that understands the physical world of flesh and bone. Corpus pairs that long-term promise with immediate practical value, even before AGI arrives: Drug developers have more promising ideas than they can test because clinical trials—the proving ground for a medicine’s safety and effectiveness in people—are expensive, limited in number by the availability of eligible patients, and slow because disease outcomes must unfold before they can be measured. A candidate that fails late can consume years of work and capital that could have supported another promising medicine, making earlier warning of potential safety problems especially valuable.
Corpus helps researchers investigate those risks by showing where a protein targeted in diseased tissue—or another protein similar enough for the medicine to bind to it unintentionally—also appears in healthy tissues, highlighting potential sites of unintended harm. Developers can use that information in existing workflows to make better-informed decisions about how to allocate capital and scarce clinical trial capacity.

As a demonstration of this immediate value, Ambitious analyzed healthy tissue expression for 973 proteins (Figure 2) being pursued as drug targets, but not yet addressed by an approved medicine. These targets merit particular scrutiny, because no medicine aimed at them has yet reached broad real-world use where target-mediated side effects can alert developers to healthy-tissue expression that existing references missed. For 19% of these targets, HPA detected no healthy expression at all, while Corpus reported it for the first time. For 77%, Corpus significantly expands the list of healthy tissues in which expression is detected. Seeing these targets in more parts of the healthy body gives developers a wider field in which to investigate risk before committing to further development.

Whereas Corpus begins with healthy human proteomics, Ambitious is already extending the same system across the continuum from health to disease, additional molecular modalities and other organisms, while increasing both the resolution and depth of measurement. Each expansion broadens the decisions the data can inform and the biological concepts models can learn.
Global Stakes & The Greatest Advantage
The existing field has trained on a patchwork of samples assembled by accident, using data exhaust from existing clinical and commercial workflows. Companies and countries that create data collection infrastructure from first principles will be positioned to lead biological AI, and best translate it into not only better prevention, diagnostics, and medicines, but also new industries and ecosystems.
“The lesson Apple demonstrated most clearly is that the most consequential products become foundations for entirely new industries. Ambitious is building that foundation as AI moves beyond language into biology, where there is no internet of reliable data waiting to be scraped; it has to be generated. China is building that capacity at scale, and any nation that wants to compete in biological AI will need to do the same,” said Jeff Martin, an Ambitious board member, who helped drive Apple’s growth in China as a member of Steve Jobs’ executive team.
“The power of Ambitious is not data, but the machine we built to generate what biological intelligence needs next,” said Hudson. “As the frontier becomes more competitive, durable advantages will increasingly come from the proprietary evidence used to train, adapt and evaluate them. Better biological measurements enable better models. Better models more efficiently select more informative hypotheses and experiments. Those experiments beget still more valuable data.”