1. From sequence to product, in one auditable loop
Evolrix AI is an AI-driven biomanufacturing platform — or more precisely, an AI-native DBTL (Design–Build–Test–Learn) loop[8]. The long-term vision is for Kairos, its AI-scientist agent, to orchestrate the full chain — sequence design, expression optimization, fermentation, and purification. Today, what is delivered is the dry-lab design half — sequence design, expression-construct optimization, and structure-confidence checks — while fermentation, scale-up, and purification are on the roadmap, reached through wet-lab partners and the active-learning loop. Every step is traceable to its evidence. The platform is being built host-by-host. Pichia comes first, a host with FDA-approved products and a deep regulatory track record; more hosts follow on the roadmap. Each capability is grounded in cited evidence and kept separate from model prediction, so that what is known, what is inferred, and what is designed can always be inspected.
By auditable loop we mean something concrete, not a slogan. Every output the platform produces — a designed sequence, a recommended construct, a structure prediction, a fermentation parameter — carries its full provenance: the original user query, the tools that were called, the arguments passed to each tool, the retrieved literature and patents, and the line of reasoning that connected them. A reviewer can reopen any output and walk back to the exact source records and the model's stated confidence. Nothing in the loop is asserted without an attached chain of evidence; what is measured is labeled as measured, what is predicted is labeled as predicted.
The DBTL framing is not borrowed decoration — it states where Evolrix AI puts its weight. Most DBTL platforms accelerate the Build/Test hardware (the foundry); Evolrix AI accelerates the Design and Learn phases, using AI plus verifiable tools plus a unique patent/Chinese-data asset to turn the most manual, experience-driven step of Pichia expression (choosing the construct and process) into a data-driven, auditable recommendation[8].
The platform rests on three data pillars, each playing a distinct role in the loop:
| Pillar | What it is | Role in the loop |
|---|---|---|
| Mechanistic knowledge | 66k full-text corpus (1,950 Pichia-focused) distilled into a knowledge graph | Retrieval & grounding for every query |
| Structure | Boltz-2 structure verification | Predicted structure-confidence screen — flags low-confidence or structurally suspicious candidates before build (in-silico, not experimental validation) |
| Dynamic expression | Host-engineering dataset, ~627 regulatory records across 21 cross-source folds[5] | Learns construct effects to give directional expression ranking |
This positioning is not abstract. It sits inside a global technology race that has hardened into national policy across every major economy — a race whose decisive lever, named independently by nearly every national strategy, is exactly the one Evolrix AI has chosen (see §2).
2. Why now: a global technology race
What justifies a dedicated platform for AI-driven biomanufacturing right now is not a single breakthrough but a convergence: every major economy has, independently, elevated synthetic biology and biomanufacturing to a national-strategic — even "tech-sovereignty" — priority, and nearly every national strategy names the same levers: AI + DBTL + data + biofoundry[14]. This is not background color; it is the demand curve under Evolrix AI's market.
Global policy: a tech-sovereignty race
| Jurisdiction | Move | What it names |
|---|---|---|
| United States | Executive Order 14081 (2022) — National Biotechnology & Biomanufacturing Initiative; CHIPS & Science Act; DoD BioMADE institute | Whole-of-government biomanufacturing; "global industry is on the cusp of an industrial revolution powered by biotechnology" |
| United Kingdom | National Vision for Engineering Biology (2023) — one of five critical technologies, £2B over 10 years | European leader; market projected at $3B by 2030 |
| EU / Germany | March-2024 measures to accelerate biotech & biomanufacturing; 12 European countries with dedicated bioeconomy strategies | Acceleration of the biomanufacturing base |
| South Korea | World's first dedicated Synthetic Biology Promotion Act (2025, effective 2026); designated a National Strategic Technology | 90% tech parity with the US by 2030; 30% of manufacturing shifted to bio within a decade; government-built public biofoundries |
| Japan | Bioeconomy Strategy | ¥100 trillion market target |
Economic scale: a multi-trillion-dollar reshaping of manufacturing
The UK government cites engineering biology producing $2–4 trillion per year of global economic impact over 2030–2040; some estimates put biomanufacturing's value by the end of the century at ~$30 trillion, on the order of one-third of global manufacturing[14]. This is the macro frame inside which the synthetic-biology market sits: US$19.75B (2025) → US$56.48B (2031) at a CAGR of ≈19%, with AI integration explicitly named a core growth driver and Asia-Pacific the fastest-growing region[10].
China: a distinctive playbook, and where Evolrix AI fits
China has put biomanufacturing at the highest tier of national planning. The 15th Five-Year Plan (2026–2030) outline lists biomanufacturing as one of six prospective "future industries" to become a new engine of economic growth, a positioning set by the CPC Central Committee's Recommendations for the 15th Five-Year Plan (4th Plenum, 2025)[9]. The Ministry of Industry and Information Technology is drafting the 15th Five-Year Biomanufacturing Development Plan, explicitly calling for "flagship products and AI application cases" and the cultivation of pilot-scale platforms — Evolrix AI maps directly onto this policy priority[9]. The scale is concrete: China's biomanufacturing total was about ¥1.1 trillion in 2025 (with >70% of global fermentation capacity), projected to reach about ¥1.8 trillion by 2030 (≈25% of the global market)[9].
Why Evolrix AI aligns with this policy direction
Every national strategy above names the same decisive levers — AI, DBTL, and data — as the bottleneck to move biomanufacturing from trial-and-error to rational design. Evolrix AI is an AI-native DBTL loop[8]: Kairos, its AI-scientist agent, is designed to orchestrate the full chain from sequence design to purification, every step traceable to its evidence — with the dry-lab design phases delivered today and fermentation/purification on the roadmap. Most DBTL platforms accelerate the Build/Test hardware (the foundry); Evolrix AI accelerates the Design and Learn phases, using AI plus verifiable tools plus a unique patent/Chinese-data asset to turn the most manual, experience-driven step of Pichia expression (choosing the construct and process) into a data-driven, auditable recommendation. Evolrix AI is betting on this exact leverage — AI × DBTL × data — landed on a core chassis organism (Pichia) with a unique data moat.
3. Where Evolrix AI sits
No virtual-cell flagship works on industrial protein production. Pichia AI is fragmented: most tooling targets general protein engineering without an industrial-host focus. The landscape splits cleanly into three camps, and Evolrix AI occupies the third.
| Player | Focus | Niche |
|---|---|---|
| Arc Institute / CZI / GenBio | Virtual cell | Human / disease cells, drug discovery |
| Cradle / Basecamp | AI protein engineering | Pharma + chemicals, multi-host |
| Evolrix AI | Pichia biomanufacturing | Full-chain recombinant-protein production, patent + Chinese data moat |
Direct peers in AI protein engineering include Cradle and Basecamp. Evolrix AI differs in two concrete ways — the Pichia host, whose products are FDA-approved and industrially mature, and its data moat drawn from patents and the Chinese-language literature, a source most English-centric tools do not mine. The first difference is a host choice; the second is a data-access choice. Both compound: a Pichia-focused platform trained on sources the others cannot see learns a different expression landscape than a general protein-engineering tool trained on English literature alone.
The upstream arena is large and growing fast. The global synthetic-biology market was about US$19.75B in 2025, projected to reach about US$56.48B by 2031 at a CAGR of ≈19%, with AI integration explicitly named a core growth driver and Asia-Pacific the fastest-growing region[10]. Protein expression is a subset of this arena; Evolrix AI's serviceable segment is the Pichia/yeast-host × expression-optimization slice within it.
4. Pichia = a mature host with FDA-approved products
Pichia is not an experimental host. It is a mature production host whose recombinant products include FDA-approved biologics, with a deep industrial and regulatory track record. The reasons it became a workhorse are structural: as a eukaryote it performs post-translational modifications (disulfide bonds, glycosylation, folding) that E. coli cannot; it secretes heterologous protein into the broth, simplifying downstream purification; it grows to very high cell density (>100 g/L dry cell weight) in fed-batch fermentation; and its glycosylation can be engineered — the GlycoFi/GlycoSwitch lineage humanized Pichia glycosylation, opening a path toward therapeutic glycoproteins[2]. These properties are why Pichia, not a faster-growing bacterium, is the host behind approved biologics.
- >5,000 recombinant proteins expressed.[2]
- >70 clinical or marketed products.
- >300 licensed industrial processes, and >300 biotech/pharma companies licensed to use the system.
- 3 FDA-approved biologics, produced under filed BLAs.
Filed FDA BLAs include:
| Product | BLA |
|---|---|
| Kalbitor | BLA 125277[1] |
| Jetrea | BLA 125422 |
| Semglee | BLA 761201 |
Beyond approved biologics, Pichia-expressed phospholipase C holds GRAS status, and the phytase market — a feed-enzyme workhorse produced in Pichia — is on the order of $350M[3]. Phytase is instructive as a case study: it is a heat-stable feed enzyme added to poultry and swine diets to release phosphate from phytate, and Pichia is one of its dominant production hosts precisely because it secretes the enzyme at high titer and performs the disulfide-rich folding the enzyme requires. Notably, non-methanol phytase production has already reached 20 g/L, showing the methanol-free route is industrially viable today. This is the foundation Evolrix AI builds on — not a host we hope will work, but one whose products are already on shelves and in feeds.
5. Two validated capabilities
Evolrix AI has two capabilities that are already validated against data, not merely promised. The first is live; the second underpins the expression-optimization layer.
5a. Structure verification
Using Boltz-2, the platform predicts and scores the structure of a candidate sequence. The validation is a controlled contrast: a properly grounded wild-type sequence reaches a median pLDDT of 0.93, while a deliberately scrambled (ungrounded) baseline of the same composition collapses to 0.48[4]. The gap is not a benchmark number cherry-picked for marketing — it is the difference between a sequence the model can fold confidently and one it cannot. Practically, this means the platform can flag a designed sequence whose structure is poorly supported before any wet-lab work is spent on it, and it does so on the public /science/ page.
| Sequence | Median pLDDT | Interpretation |
|---|---|---|
| Wild-type (grounded) | 0.93 | High-confidence fold |
| Scrambled baseline | 0.48 | Model cannot resolve fold |
5b. Expression-construct optimization
The second capability learns which construct choices (promoter, secretion signal, chaperone co-expression, fermentation strategy) move expression, from the host-engineering dataset. The dataset has grown substantially since v0.1: it now holds ~627 regulatory records from 83 independent sources, covering 21 cross-source protein folds[5]. This capability serves two roles: (A) an evidence engine (primary, robust) — structuring which constructs have been used, what titer they yielded, and from which patent or paper; and (B) early directional ranking (secondary, honestly weak) — ranking candidate constructs by predicted relative expression. The field's data is noisy by default — across all folds, cross-source construct scores correlate only weakly — which is exactly why naive transfer of one lab's construct recipe to another lab's protein usually disappoints. Evolrix AI's response to that noise is not to hide it but to distinguish what can be computed from what still requires wet-lab, and to report the confidence on the computable axis.
The most robust evidence is method-independent: the model recovers element effects that match known biology. The learned promoter pCS1 effect (+0.7, n=62) corresponds to Lonza's commercial strong promoter; chaperones ERO1 (oxidative folding) and SBH1 (signal-peptide processing) emerge as positive contributors, consistent with Pichia folding and glycosylation mechanisms[5]. This evidence is not affected by the evaluation metric and is the most honest validation of the capability. Cross-protein generalization is method-dependent and reported as an honest range: within-source (de-confounded) ρ≈0.12 as the conservative floor; cross-source (including lab-scale effects) ρ≈0.26; overall ρ≈0.37 (dominated by single-source folds). On high-confidence folds, xylanase and mannanase titers reach ρ 0.76–0.89. These numbers are weak but positive — a genuine early directional signal, not a strong predictor. The path from here to a predictor runs through controlled wet-lab data (same protein, vary only the construct, same conditions), which is the active-learning loop on the roadmap.
| Construct choice / metric | Learned effect | Basis |
|---|---|---|
| Promoter pCS1 | +0.7 (positive) | n=62 held-out |
| Chaperone ERO1 | Positive | Recovered from data |
| Within-source ρ (de-confounded) | ≈0.12 | Conservative floor |
| Cross-source ρ (with lab-scale) | ≈0.26 | 21 cross-source folds, ~627 records |
| High-confidence folds (xylanase/mannanase) | ρ 0.76–0.89 | Cross-source titer |
What this means practically: instead of treating expression as a single lab's recipe to be copied, the platform quantifies how much each construct lever contributes, and tells the user the confidence (the ρ and the effect size) rather than a bare titer prediction. The platform distinguishes the computable part of the problem (which construct levers move expression, and by how much, given cross-source data) from the part that still requires wet-lab validation (the absolute titer for a specific protein in a specific lab) — and it is honest about which is which. The output is directional evidence for construct design, not a guarantee of titer, and that distinction is itself a feature for any customer who needs to justify a build decision with an auditable chain.
Dataset & evaluation snapshot
The single source of truth for these numbers is the shared site config (site-truth.js); this section mirrors it for readability. Last verified 2026-07-12.
| Metric | Value | Note |
|---|---|---|
| Snapshot date | 2026-07-06 | host-eng pipeline v0.4 |
| Records (total) | 664 | 627 regulatory + 28 glyco + 23 lit + 9 seq |
| Cross-source folds | 24 total (21 regulatory-only cohort) | These two numbers are not contradictory — 24 includes glyco/lit/seq folds; 21 is the regulatory-only cohort used for the primary ρ figures |
| Sources | 83 independent | Patents + literature |
| Within-source ρ (de-confounded) | ≈0.12 | Conservative floor — cross-protein construct ranking |
| Cross-source ρ | ≈0.26 | Includes lab-scale effects |
| Overall ρ | ≈0.37 | Dominated by single-source folds |
| High-confidence folds | ρ 0.76–0.89 | Xylanase / mannanase (larger samples) |
| Validation status | Early directional signal — not a titer predictor. Cross-source noise; small-sample folds (lipase, insulin) still weak; absolute titer requires wet-lab. | |
Important metric distinction: The homepage figure host-fidelity ρ≈0.35 is a different metric — it measures within-protein codon-design reward correlation with real titers (how well the CDS design ranks within one protein), not the cross-protein construct-expression ranking (ρ≈0.12–0.37) described above. Do not conflate the two.
6. Data moat: patents + Chinese
Evolrix AI mines patents and the Chinese-language literature, two sources that most English-centric protein-engineering tools ignore. This is where much of the practical, industrially relevant Pichia know-how actually lives — patent filings disclose construct details, titers, and process conditions that journal papers often omit, and a large share of the world's industrial-enzyme and feed-enzyme Pichia work is published in Chinese. The data assets stack into three layers:
| Layer | Scale | Use |
|---|---|---|
| Literature corpus | 66k full-text (1,950 Pichia) + knowledge graph | Retrieval / grounding |
| Host-engineering structured | 664 total (~627 regulatory + 28 glyco + 23 lit + 9 seq), 24 cross-source folds, 83 sources | Construct optimization |
| Lactoferrin case | 4 Chinese institutions | Cross-source validation proof |
Expansion progress on the host-engineering layer is explicit: cross-source folds went from 9 → 24, exceeding the 15+ target set in v0.1. The next steps are filling weak small-sample folds (lipase, insulin) and cleaning the regulatory-element vocabulary to ship the precise recommender. The evidence is concrete: for lactoferrin, the platform draws expression data from four Chinese institutions — ZenoBio, Jiangnan University, Mengniu, and Sanyuan[7] — a corpus no English-only retriever can reach. The point of the lactoferrin case is not the protein itself but the proof of coverage: if the platform can cross-validate one protein across four independent Chinese sources, the same retrieval reaches the rest of the Chinese-language Pichia corpus in the same way.
Coverage ledger & honest limits
Authoritative corpus counts (verified 2026-07-12 via backend corpus_counts.py):
| Source layer | Records | Distinct patents | Curation | Last updated | Known gaps |
|---|---|---|---|---|---|
| Patent + literature evidence (all CN) | 2,580 | 386 | Mixed: 186 curated high / 785 medium / 1,612 needs-verify | 2026-07-06 | CNIPA/CNKI full-text walls block API access — largest gap |
| Host-engineering structured | 664 | 83 sources | Curated (regulatory/glyco/lit/seq) | 2026-07-06 | Small-sample folds weak (lipase, insulin) |
| Literature corpus (full-text) | 66k (1,950 Pichia) | — | Qdrant indexed | 2026-06 | Sequence-variant data effectively absent (regime mismatch) |
Honest framing: This is not a "complete Chinese-language Pichia moat." It is a high-quality but partial corpus — 386 distinct patents and 2,580 evidence records, heavily CN-centric, with real coverage gaps (CNIPA/CNKI full-text access walls, weak small-sample folds). The claim is that the method (retrieval + structured extraction + cross-source validation) is sound and the lactoferrin 4-institution case proves the approach; the claim is not that coverage is exhaustive. We state this plainly because a customer building a regulated decision needs to know exactly where the evidence thins.
7. Non-methanol expression systems
Methanol-inducible AOX1 is canonical for Pichia and gives very high titers, but industrial production increasingly needs non-methanol systems — for safety, scale, and regulatory reasons. This is an industrial need, not a research curiosity. Methanol is flammable and acutely toxic — lethal to the cells themselves at 2–5% v/v[11] — carries a high heat of combustion (high oxygen demand), and triggers stricter fire-safety classification, explosion-proof infrastructure, and strong cooling for large fermenters. Methanol-free regimes remove that overhead, and the industry is steadily migrating to them. Evolrix AI's host-engineering data covers the full non-methanol promoter toolbox:
| Promoter | Regulation | Use case |
|---|---|---|
| pGAP | Constitutive (glyceraldehyde-3-phosphate dehydrogenase) | Most common constitutive; near AOX1 level, no inducer |
| pGCW14 / pUPP | Constitutive (commercial GCW14 variant) | Strong constitutive; pUPP/PDF reach up to 9× pGAP for CalB[12] |
| pTEF1 | Constitutive (translation elongation factor) | High constitutive expression |
| pPGK1 | Constitutive (phosphoglycerate kinase) | Steady constitutive expression |
| PDH / PDF | Derepressed inducible (feed-rate controlled) | Both "inducible" and "methanol-free"; controlled by feed rate alone[12] |
Why methanol-free matters in practice: it removes a flammable-toxic inducer from the process, simplifies scale-up (no methanol feed strategy or off-gas handling), and eases regulatory filings where methanol classification is a burden. The literature shows PUPP/PDF reaching up to 9× the specific productivity of pGAP for CalB[12] — meaning "choosing the right methanol-free system" itself has large optimization headroom, exactly where the model's recommendation adds value. The platform's data already covers both methanol and methanol-free constructs (e.g. the learned pCS1 and pGAP effects), so it can directly answer "methanol-induced versus methanol-free constitutive, and which promoter," and give scenario-specific recommendations by product type, methanol-free requirement, and scale-up stage.
8. Product boundary
Evolrix AI's product boundary is recombinant proteins plus active molecules. The protein side spans industrial enzymes and therapeutic proteins; the active-molecule side extends the host into high-value small molecules via metabolic engineering. Concrete examples beyond the flagship case:
| Category | Examples | Why Pichia |
|---|---|---|
| Industrial enzymes | Phytase, lipase, cellulase | High-density secretion, disulfide folding |
| Therapeutic proteins | HSA, insulin | Eukaryotic PTM, GRAS-track host |
| Active small molecules | α-santalene (21.5 g/L), hyaluronic acid (0.8–1.7 g/L), 3-HP | Mevalonate / metabolic pathway engineering[6] |
The active-molecule side is not aspirational: α-santalene has been produced at 21.5 g/L[6] in Pichia via an engineered mevalonate pathway, hyaluronic acid at 0.8–1.7 g/L, and 3-HP from methanol — showing the host's reach beyond proteins into high-value small molecules. Industrial enzymes like phytase, lipase, and cellulase are the bread-and-butter of Pichia manufacturing today; therapeutic proteins like human serum albumin (HSA) and insulin exploit the host's eukaryotic folding and secretion. The platform's scope is the union of these, not a single product class.
9. Commercialization & landing
Evolrix AI's commercial path targets three customer types, defined by who actually runs Pichia protein expression:
| Customer | Pain point | Entry |
|---|---|---|
| Research institutes — university/institute synthetic-biology and protein-engineering labs | Construct selection by trial-and-error; time-consuming literature search | Low-barrier tool + auditable evidence; build reputation and a data flywheel first |
| Pharmaceutical companies — biologics / recombinant-protein R&D teams (incl. CROs/CDMOs) | Slow expression optimization; high screening cost | Accelerate expression-construct design, shortening the gene-to-first-product cycle |
| Synthetic-biology companies — industrial-enzyme / protein / biomanufacturing firms | Raising titer; lowering fermentation cost | Construct/process recommendation + structure verification, owning titer and cost directly |
The market these customers sit in is large. As above, the upstream synthetic-biology market is ~US$19.75B (2025) → ~US$56.48B (2031) at CAGR ≈19%[10]; the more direct protein-expression market is ~US$3.0–5.1B (2025) → ~US$5.0–7.7B (2031/2035) at CAGR ≈8–9%, with CROs/CDMOs the fastest-growing end-user and Asia-Pacific the fastest-growing region. A benchmark exists for the model: Cradle Bio serves exactly pharma (Novo Nordisk, J&J) plus chemicals/food/agriculture, confirming the reality of "research + pharma + synth-bio" customers; Evolrix AI's difference is a focus on the Pichia host and Chinese/patent data.
Referencing the validated Cradle Bio business model, a three-tier pricing structure is proposed:
| Tier | Form | For |
|---|---|---|
| Entry / academic | Low-cost or free web tool (evolrix.bio/science) | Research institutes — reputation + data flywheel |
| Pro / API | Subscription + full API into customer pipelines; customer data trains a private model, IP stays with the customer | Pharma / synth-bio R&D teams |
| Enterprise / project | Per-project / per-molecule + private deployment | Large pharma / CDMO — accelerating R&D as a quantifiable value anchor |
The quantifiable value propositions anchor these tiers: time saved (from "search literature + trial-and-error construct selection" to "model recommendation + evidence provenance"); higher titer / lower cost (construct/process recommendations own expression directly); trustworthy / auditable (every recommendation carries patent/literature provenance + structure verification, distinct from black-box tools); and a data moat → customer flywheel (customer wet-lab results feed back and continuously improve the model, on Evolrix AI's unique Chinese/patent data base). Specific pricing figures will be calibrated with early customers, as AI protein platforms are mostly enterprise/quote-based and Cradle/Basecamp do not publish unit prices.
10. Intellectual property
Evolrix AI is building an IP portfolio across software copyrights and invention patents, with filings in progress:
- Software copyrights (~12 planned), covering: the Kairos agent orchestration layer, the host-engineering calibration engine, the structure-verification tool, and the retrieval/knowledge-graph modules.
- Invention patents (~6–8 planned), with directions including: (1) a discrete-regulatory-feature method for expression-construct recommendation; (2) a verifiable AI research-workbench architecture; (3) a structured-mining-and-deconfounding pipeline for patent/literature expression data.
- Competitions / credentials: participating in the National Disruptive Technology Innovation Competition.
Two boundaries should be stated plainly alongside the commercial and IP claims — and we treat stating them plainly as a trust credential, not a limitation. A platform that can tell you precisely where its confidence ends is a platform you can build a regulated decision on. First, Boltz-2 structure outputs are predictions, not experimental structures — they are grounded and well-calibrated, but they are not crystallographic or cryo-EM data; the platform labels them as predictions and never presents them as measured. Second, expression-construct effects are directional evidence from cross-source correlation[5], not guaranteed outcomes — the platform distinguishes the computable axis (within-source ρ≈0.12, cross-source ρ≈0.26) from the wet-lab axis (absolute titer for a specific case), and flags the latter as requiring validation before any production claim. Claims are kept to what the evidence supports; predictions are labeled as predictions, and construct effects as directional evidence requiring wet-lab validation. In an industry where black-box "it will work" promises are the norm, this discipline is the basis on which a pharma or CDMO customer can actually adopt the output.
11. Roadmap
The full chain is sequence design → expression → fermentation → purification. The state of each layer is given as a capability map — what is validated today, what is in progress, what is planned — so that an investor or partner can see exactly where the platform's evidence sits and where the next inflection points are:
Design layer — validated
- Kairos agent orchestration (reason → retrieve → design → verify → artifact).
- Structure verification via Boltz-2, live on /science/ (measured 0.93 / 0.48).
- Expression-construct optimization working, cross-source ρ stable; fermentation-condition dimension included.
- Whitepaper v0.5 (this document): evidence thickening with 24 cross-source folds.
Manufacturing layer — planned, needs wet-lab partners
- Bench → pilot scale-up (fermentation engineering): planned, requires process data and wet-lab partners.
- Downstream separation / purification (DSP): planned, yield/purity optimization.
- Data-feedback loop: partial within the design layer; manufacturing feedback to be built.
Near & mid term
- Precise recommender online (regulatory-element vocabulary cleanup).
- Active-learning loop — pick the most informative wet-lab experiments (the Learn phase of DBTL).
- Kairos completion (reviewer + retrieval into the agent).
- Wet-lab / process partners → open the manufacturing layer (bench → pilot → purification).
- Multi-host expansion — transfer the methodology from Pichia to E. coli, other yeasts, and eventually mammalian hosts once validated.
- Whitepaper v1.0 (commercial figures + IP filing numbers).
The two layers are not separate pipelines; they are connected by an active-learning spine. As the manufacturing layer runs — real fermentation batches with measured titers and real purification yields — that outcome data flows back into the host-engineering dataset. A construct that over- or under-performs the platform's prediction becomes a new labeled record, sharpening the next ρ estimate and the next effect-size call. The design layer is validated today; the manufacturing layer is on the roadmap, connected to the same evidence-and-provenance spine so that the loop stays auditable end to end.
References
- FDA Center for Drug Evaluation and Research. Biologics License Applications (BLA) and approvals: Kalbitor (BLA 125277), Jetrea (BLA 125422), Semglee (BLA 761201). U.S. Food and Drug Administration. FDA Biologics Approvals.
- Zha J, et al. Advances in Metabolic Engineering of Pichia pastoris Strains as Powerful Cell Factories. J Fungi (Basel), 2023. doi:10.3390/jof9101027.
- Market data: phytase market size (~$350M/yr) and Pichia commercialization scale (>5,000 proteins, >70 products, >300 licensed industrial processes, >300 biotech/pharma companies). Mordor Intelligence / MarketsandMarkets, 2025–2026 reports; Biomolecules 13(3):441 (2023); Front. Biosci.-Elite (2024).
- Evolrix AI Science. Structure verification live demo: ubiquitin wild-type pLDDT 0.93 vs scrambled 0.48 (Boltz-2 prediction). evolrix.bio/science.
- Evolrix AI host-engineering dataset (v0.5): ~627 regulatory records from 83 independent sources; 21 cross-source folds; within-source (de-confounded) Spearman ρ≈0.12; cross-source ρ≈0.26; overall ρ≈0.37; high-confidence folds xylanase/mannanase ρ 0.76–0.89; learned promoter pCS1 effect +0.7 (n=62); chaperones ERO1/SBH1 positive. evolrix.bio/research.
- α-Santalene production in Pichia pastoris at reported titer 21.5 g/L; hyaluronic acid 0.8–1.7 g/L; 3-HP from methanol. Zha J, et al. J Fungi 9(10):1027 (2023); Metabolic engineering of K. phaffii for 3-HP production from methanol, PubMed 39979934 (2025).
- Human lactoferrin expression data from 4 independent Chinese institutions: CN118147180A (ZenoBio), CN115960175A (Jiangnan University + Mengniu), CN116970503B (Jiangnan University), CN1718726A (Sanyuan). Chinese patent database (CNIPA / Google Patents).
- NCBI Bookshelf. Design-Build-Test-Learn: Impact of AI on the Synthetic Biology Process (The Age of AI in the Life Sciences) — DBTL canonical cycle; AI/ML entering the Design and Learn phases, foundries automating Build/Test.
- 15th Five-Year Plan (2026–2030) sources: CPC Central Committee Recommendations for the 15th Five-Year Plan (4th Plenum, 2025) + 15th Five-Year Plan outline (biomanufacturing as one of six prospective future industries, NDRC/Xinhua 2026-04); MIIT drafting the 15th Five-Year Biomanufacturing Development Plan (Xinhua 2025-12-19: flagship products + AI application cases, pilot-platform cultivation); 2025 Biomanufacturing Conference data (total ~¥1.1 trillion, >70% of global fermentation capacity); iCapital forecast 2030 China ~¥1.8 trillion, ~25% of global.
- Mordor Intelligence. Synthetic Biology Market 2026–2031 — US$19.75B (2025) → US$56.48B (2031), CAGR 19.14%; genome engineering 33.21% share, Asia-Pacific fastest-growing; AI integration named a core growth driver.
- Bi-directionalized promoter systems for methanol-free peroxygenases, Microb. Cell Fact. (2024) — methanol toxicity (2–5% v/v lethal to cells); Engineering the expression system for K. phaffii: a methanol-free expression system, PMC6736287 — GAP/DAS1 engineering, methanol-free routes.
- Bioprocess performance of novel methanol-independent promoters (PDF/PUPP), Microb. Cell Fact. (2021) — PUPP/PDF up to 9× PGAP for CalB; Enabling growth-decoupled K. phaffii production based on the methanol-free PDH promoter, PMC10076887.
- Jacobs P, et al. Engineering complex-type N-glycosylation in Pichia using GlycoSwitch. Nat. Protoc. 4:58–70 (2009); pichia.com GlycoSwitch — OCH1 knockout, humanized glycosylation, SuperMan5.
- Global synthetic-biology national strategies [R43]: US Executive Order 14081 (2022-09) + CHIPS & Science Act National Engineering Biology R&D Initiative + DoD BioMADE; UK National Vision for Engineering Biology (DSIT, 2023, £2B/10yr, one of five critical technologies); EU March-2024 biotech/biomanufacturing acceleration measures + EC bioeconomy country strategies (JRC, 2025-09); South Korea Synthetic Biology Promotion Act (passed 2025, effective 2026, world's first dedicated law, MSIT) + National Strategic Technology (2023-12); Japan Bioeconomy Strategy (¥100T market target). Economic scale: McKinsey The Bio Revolution and UK business.gov.uk engineering biology — $2–4T/yr global impact (2030–2040), ~$30T by end of century (~1/3 of global manufacturing). Reviews: SynBioBeta The Influence of Synthetic Biology on National Bioeconomy Strategies (2024), OECD synthetic-biology governance report.