← Research

Whitepaper

The Evolrix AI Platform Whitepaper

July 4, 2026

From sequence to product, in one auditable loop — an AI-driven biomanufacturing platform where the Kairos AI-scientist agent orchestrates the full chain, with every step traceable to its evidence.

1. From sequence to product, in one auditable loop

Evolrix AI is an AI-driven biomanufacturing platform — or more precisely, an AI-native DBTL (Design–Build–Test–Learn) loop[8]. The long-term vision is for Kairos, its AI-scientist agent, to orchestrate the full chain — sequence design, expression optimization, fermentation, and purification. Today, what is delivered is the dry-lab design half — sequence design, expression-construct optimization, and structure-confidence checks — while fermentation, scale-up, and purification are on the roadmap, reached through wet-lab partners and the active-learning loop. Every step is traceable to its evidence. The platform is being built host-by-host. Pichia comes first, a host with FDA-approved products and a deep regulatory track record; more hosts follow on the roadmap. Each capability is grounded in cited evidence and kept separate from model prediction, so that what is known, what is inferred, and what is designed can always be inspected.

By auditable loop we mean something concrete, not a slogan. Every output the platform produces — a designed sequence, a recommended construct, a structure prediction, a fermentation parameter — carries its full provenance: the original user query, the tools that were called, the arguments passed to each tool, the retrieved literature and patents, and the line of reasoning that connected them. A reviewer can reopen any output and walk back to the exact source records and the model's stated confidence. Nothing in the loop is asserted without an attached chain of evidence; what is measured is labeled as measured, what is predicted is labeled as predicted.

The DBTL framing is not borrowed decoration — it states where Evolrix AI puts its weight. Most DBTL platforms accelerate the Build/Test hardware (the foundry); Evolrix AI accelerates the Design and Learn phases, using AI plus verifiable tools plus a unique patent/Chinese-data asset to turn the most manual, experience-driven step of Pichia expression (choosing the construct and process) into a data-driven, auditable recommendation[8].

Evolrix full-chain loop
The full-chain loop: sequence design → expression optimization → fermentation → purification, orchestrated by Kairos with evidence attached at every step.

The platform rests on three data pillars, each playing a distinct role in the loop:

PillarWhat it isRole in the loop
Mechanistic knowledge66k full-text corpus (1,950 Pichia-focused) distilled into a knowledge graphRetrieval & grounding for every query
StructureBoltz-2 structure verificationPredicted structure-confidence screen — flags low-confidence or structurally suspicious candidates before build (in-silico, not experimental validation)
Dynamic expressionHost-engineering dataset, ~627 regulatory records across 21 cross-source folds[5]Learns construct effects to give directional expression ranking

This positioning is not abstract. It sits inside a global technology race that has hardened into national policy across every major economy — a race whose decisive lever, named independently by nearly every national strategy, is exactly the one Evolrix AI has chosen (see §2).

2. Why now: a global technology race

What justifies a dedicated platform for AI-driven biomanufacturing right now is not a single breakthrough but a convergence: every major economy has, independently, elevated synthetic biology and biomanufacturing to a national-strategic — even "tech-sovereignty" — priority, and nearly every national strategy names the same levers: AI + DBTL + data + biofoundry[14]. This is not background color; it is the demand curve under Evolrix AI's market.

Global policy: a tech-sovereignty race

JurisdictionMoveWhat it names
United StatesExecutive Order 14081 (2022) — National Biotechnology & Biomanufacturing Initiative; CHIPS & Science Act; DoD BioMADE instituteWhole-of-government biomanufacturing; "global industry is on the cusp of an industrial revolution powered by biotechnology"
United KingdomNational Vision for Engineering Biology (2023) — one of five critical technologies, £2B over 10 yearsEuropean leader; market projected at $3B by 2030
EU / GermanyMarch-2024 measures to accelerate biotech & biomanufacturing; 12 European countries with dedicated bioeconomy strategiesAcceleration of the biomanufacturing base
South KoreaWorld's first dedicated Synthetic Biology Promotion Act (2025, effective 2026); designated a National Strategic Technology90% tech parity with the US by 2030; 30% of manufacturing shifted to bio within a decade; government-built public biofoundries
JapanBioeconomy Strategy¥100 trillion market target

Economic scale: a multi-trillion-dollar reshaping of manufacturing

The UK government cites engineering biology producing $2–4 trillion per year of global economic impact over 2030–2040; some estimates put biomanufacturing's value by the end of the century at ~$30 trillion, on the order of one-third of global manufacturing[14]. This is the macro frame inside which the synthetic-biology market sits: US$19.75B (2025) → US$56.48B (2031) at a CAGR of ≈19%, with AI integration explicitly named a core growth driver and Asia-Pacific the fastest-growing region[10].

China: a distinctive playbook, and where Evolrix AI fits

China has put biomanufacturing at the highest tier of national planning. The 15th Five-Year Plan (2026–2030) outline lists biomanufacturing as one of six prospective "future industries" to become a new engine of economic growth, a positioning set by the CPC Central Committee's Recommendations for the 15th Five-Year Plan (4th Plenum, 2025)[9]. The Ministry of Industry and Information Technology is drafting the 15th Five-Year Biomanufacturing Development Plan, explicitly calling for "flagship products and AI application cases" and the cultivation of pilot-scale platforms — Evolrix AI maps directly onto this policy priority[9]. The scale is concrete: China's biomanufacturing total was about ¥1.1 trillion in 2025 (with >70% of global fermentation capacity), projected to reach about ¥1.8 trillion by 2030 (≈25% of the global market)[9].

Why Evolrix AI aligns with this policy direction

Every national strategy above names the same decisive levers — AI, DBTL, and data — as the bottleneck to move biomanufacturing from trial-and-error to rational design. Evolrix AI is an AI-native DBTL loop[8]: Kairos, its AI-scientist agent, is designed to orchestrate the full chain from sequence design to purification, every step traceable to its evidence — with the dry-lab design phases delivered today and fermentation/purification on the roadmap. Most DBTL platforms accelerate the Build/Test hardware (the foundry); Evolrix AI accelerates the Design and Learn phases, using AI plus verifiable tools plus a unique patent/Chinese-data asset to turn the most manual, experience-driven step of Pichia expression (choosing the construct and process) into a data-driven, auditable recommendation. Evolrix AI is betting on this exact leverage — AI × DBTL × data — landed on a core chassis organism (Pichia) with a unique data moat.

3. Where Evolrix AI sits

No virtual-cell flagship works on industrial protein production. Pichia AI is fragmented: most tooling targets general protein engineering without an industrial-host focus. The landscape splits cleanly into three camps, and Evolrix AI occupies the third.

PlayerFocusNiche
Arc Institute / CZI / GenBioVirtual cellHuman / disease cells, drug discovery
Cradle / BasecampAI protein engineeringPharma + chemicals, multi-host
Evolrix AIPichia biomanufacturingFull-chain recombinant-protein production, patent + Chinese data moat

Direct peers in AI protein engineering include Cradle and Basecamp. Evolrix AI differs in two concrete ways — the Pichia host, whose products are FDA-approved and industrially mature, and its data moat drawn from patents and the Chinese-language literature, a source most English-centric tools do not mine. The first difference is a host choice; the second is a data-access choice. Both compound: a Pichia-focused platform trained on sources the others cannot see learns a different expression landscape than a general protein-engineering tool trained on English literature alone.

The upstream arena is large and growing fast. The global synthetic-biology market was about US$19.75B in 2025, projected to reach about US$56.48B by 2031 at a CAGR of ≈19%, with AI integration explicitly named a core growth driver and Asia-Pacific the fastest-growing region[10]. Protein expression is a subset of this arena; Evolrix AI's serviceable segment is the Pichia/yeast-host × expression-optimization slice within it.

4. Pichia = a mature host with FDA-approved products

Pichia is not an experimental host. It is a mature production host whose recombinant products include FDA-approved biologics, with a deep industrial and regulatory track record. The reasons it became a workhorse are structural: as a eukaryote it performs post-translational modifications (disulfide bonds, glycosylation, folding) that E. coli cannot; it secretes heterologous protein into the broth, simplifying downstream purification; it grows to very high cell density (>100 g/L dry cell weight) in fed-batch fermentation; and its glycosylation can be engineered — the GlycoFi/GlycoSwitch lineage humanized Pichia glycosylation, opening a path toward therapeutic glycoproteins[2]. These properties are why Pichia, not a faster-growing bacterium, is the host behind approved biologics.

Filed FDA BLAs include:

ProductBLA
KalbitorBLA 125277[1]
JetreaBLA 125422
SemgleeBLA 761201

Beyond approved biologics, Pichia-expressed phospholipase C holds GRAS status, and the phytase market — a feed-enzyme workhorse produced in Pichia — is on the order of $350M[3]. Phytase is instructive as a case study: it is a heat-stable feed enzyme added to poultry and swine diets to release phosphate from phytate, and Pichia is one of its dominant production hosts precisely because it secretes the enzyme at high titer and performs the disulfide-rich folding the enzyme requires. Notably, non-methanol phytase production has already reached 20 g/L, showing the methanol-free route is industrially viable today. This is the foundation Evolrix AI builds on — not a host we hope will work, but one whose products are already on shelves and in feeds.

5. Two validated capabilities

Evolrix AI has two capabilities that are already validated against data, not merely promised. The first is live; the second underpins the expression-optimization layer.

5a. Structure verification

Using Boltz-2, the platform predicts and scores the structure of a candidate sequence. The validation is a controlled contrast: a properly grounded wild-type sequence reaches a median pLDDT of 0.93, while a deliberately scrambled (ungrounded) baseline of the same composition collapses to 0.48[4]. The gap is not a benchmark number cherry-picked for marketing — it is the difference between a sequence the model can fold confidently and one it cannot. Practically, this means the platform can flag a designed sequence whose structure is poorly supported before any wet-lab work is spent on it, and it does so on the public /science/ page.

SequenceMedian pLDDTInterpretation
Wild-type (grounded)0.93High-confidence fold
Scrambled baseline0.48Model cannot resolve fold

5b. Expression-construct optimization

The second capability learns which construct choices (promoter, secretion signal, chaperone co-expression, fermentation strategy) move expression, from the host-engineering dataset. The dataset has grown substantially since v0.1: it now holds ~627 regulatory records from 83 independent sources, covering 21 cross-source protein folds[5]. This capability serves two roles: (A) an evidence engine (primary, robust) — structuring which constructs have been used, what titer they yielded, and from which patent or paper; and (B) early directional ranking (secondary, honestly weak) — ranking candidate constructs by predicted relative expression. The field's data is noisy by default — across all folds, cross-source construct scores correlate only weakly — which is exactly why naive transfer of one lab's construct recipe to another lab's protein usually disappoints. Evolrix AI's response to that noise is not to hide it but to distinguish what can be computed from what still requires wet-lab, and to report the confidence on the computable axis.

The most robust evidence is method-independent: the model recovers element effects that match known biology. The learned promoter pCS1 effect (+0.7, n=62) corresponds to Lonza's commercial strong promoter; chaperones ERO1 (oxidative folding) and SBH1 (signal-peptide processing) emerge as positive contributors, consistent with Pichia folding and glycosylation mechanisms[5]. This evidence is not affected by the evaluation metric and is the most honest validation of the capability. Cross-protein generalization is method-dependent and reported as an honest range: within-source (de-confounded) ρ≈0.12 as the conservative floor; cross-source (including lab-scale effects) ρ≈0.26; overall ρ≈0.37 (dominated by single-source folds). On high-confidence folds, xylanase and mannanase titers reach ρ 0.76–0.89. These numbers are weak but positive — a genuine early directional signal, not a strong predictor. The path from here to a predictor runs through controlled wet-lab data (same protein, vary only the construct, same conditions), which is the active-learning loop on the roadmap.

Construct choice / metricLearned effectBasis
Promoter pCS1+0.7 (positive)n=62 held-out
Chaperone ERO1PositiveRecovered from data
Within-source ρ (de-confounded)≈0.12Conservative floor
Cross-source ρ (with lab-scale)≈0.2621 cross-source folds, ~627 records
High-confidence folds (xylanase/mannanase)ρ 0.76–0.89Cross-source titer

What this means practically: instead of treating expression as a single lab's recipe to be copied, the platform quantifies how much each construct lever contributes, and tells the user the confidence (the ρ and the effect size) rather than a bare titer prediction. The platform distinguishes the computable part of the problem (which construct levers move expression, and by how much, given cross-source data) from the part that still requires wet-lab validation (the absolute titer for a specific protein in a specific lab) — and it is honest about which is which. The output is directional evidence for construct design, not a guarantee of titer, and that distinction is itself a feature for any customer who needs to justify a build decision with an auditable chain.

Dataset & evaluation snapshot

The single source of truth for these numbers is the shared site config (site-truth.js); this section mirrors it for readability. Last verified 2026-07-12.

MetricValueNote
Snapshot date2026-07-06host-eng pipeline v0.4
Records (total)664627 regulatory + 28 glyco + 23 lit + 9 seq
Cross-source folds24 total (21 regulatory-only cohort)These two numbers are not contradictory — 24 includes glyco/lit/seq folds; 21 is the regulatory-only cohort used for the primary ρ figures
Sources83 independentPatents + literature
Within-source ρ (de-confounded)≈0.12Conservative floor — cross-protein construct ranking
Cross-source ρ≈0.26Includes lab-scale effects
Overall ρ≈0.37Dominated by single-source folds
High-confidence foldsρ 0.76–0.89Xylanase / mannanase (larger samples)
Validation statusEarly directional signal — not a titer predictor. Cross-source noise; small-sample folds (lipase, insulin) still weak; absolute titer requires wet-lab.

Important metric distinction: The homepage figure host-fidelity ρ≈0.35 is a different metric — it measures within-protein codon-design reward correlation with real titers (how well the CDS design ranks within one protein), not the cross-protein construct-expression ranking (ρ≈0.12–0.37) described above. Do not conflate the two.

6. Data moat: patents + Chinese

Evolrix AI mines patents and the Chinese-language literature, two sources that most English-centric protein-engineering tools ignore. This is where much of the practical, industrially relevant Pichia know-how actually lives — patent filings disclose construct details, titers, and process conditions that journal papers often omit, and a large share of the world's industrial-enzyme and feed-enzyme Pichia work is published in Chinese. The data assets stack into three layers:

LayerScaleUse
Literature corpus66k full-text (1,950 Pichia) + knowledge graphRetrieval / grounding
Host-engineering structured664 total (~627 regulatory + 28 glyco + 23 lit + 9 seq), 24 cross-source folds, 83 sourcesConstruct optimization
Lactoferrin case4 Chinese institutionsCross-source validation proof

Expansion progress on the host-engineering layer is explicit: cross-source folds went from 9 → 24, exceeding the 15+ target set in v0.1. The next steps are filling weak small-sample folds (lipase, insulin) and cleaning the regulatory-element vocabulary to ship the precise recommender. The evidence is concrete: for lactoferrin, the platform draws expression data from four Chinese institutionsZenoBio, Jiangnan University, Mengniu, and Sanyuan[7] — a corpus no English-only retriever can reach. The point of the lactoferrin case is not the protein itself but the proof of coverage: if the platform can cross-validate one protein across four independent Chinese sources, the same retrieval reaches the rest of the Chinese-language Pichia corpus in the same way.

Coverage ledger & honest limits

Authoritative corpus counts (verified 2026-07-12 via backend corpus_counts.py):

Source layerRecordsDistinct patentsCurationLast updatedKnown gaps
Patent + literature evidence (all CN)2,580386Mixed: 186 curated high / 785 medium / 1,612 needs-verify2026-07-06CNIPA/CNKI full-text walls block API access — largest gap
Host-engineering structured66483 sourcesCurated (regulatory/glyco/lit/seq)2026-07-06Small-sample folds weak (lipase, insulin)
Literature corpus (full-text)66k (1,950 Pichia)Qdrant indexed2026-06Sequence-variant data effectively absent (regime mismatch)

Honest framing: This is not a "complete Chinese-language Pichia moat." It is a high-quality but partial corpus — 386 distinct patents and 2,580 evidence records, heavily CN-centric, with real coverage gaps (CNIPA/CNKI full-text access walls, weak small-sample folds). The claim is that the method (retrieval + structured extraction + cross-source validation) is sound and the lactoferrin 4-institution case proves the approach; the claim is not that coverage is exhaustive. We state this plainly because a customer building a regulated decision needs to know exactly where the evidence thins.

7. Non-methanol expression systems

Methanol-inducible AOX1 is canonical for Pichia and gives very high titers, but industrial production increasingly needs non-methanol systems — for safety, scale, and regulatory reasons. This is an industrial need, not a research curiosity. Methanol is flammable and acutely toxic — lethal to the cells themselves at 2–5% v/v[11] — carries a high heat of combustion (high oxygen demand), and triggers stricter fire-safety classification, explosion-proof infrastructure, and strong cooling for large fermenters. Methanol-free regimes remove that overhead, and the industry is steadily migrating to them. Evolrix AI's host-engineering data covers the full non-methanol promoter toolbox:

PromoterRegulationUse case
pGAPConstitutive (glyceraldehyde-3-phosphate dehydrogenase)Most common constitutive; near AOX1 level, no inducer
pGCW14 / pUPPConstitutive (commercial GCW14 variant)Strong constitutive; pUPP/PDF reach up to 9× pGAP for CalB[12]
pTEF1Constitutive (translation elongation factor)High constitutive expression
pPGK1Constitutive (phosphoglycerate kinase)Steady constitutive expression
PDH / PDFDerepressed inducible (feed-rate controlled)Both "inducible" and "methanol-free"; controlled by feed rate alone[12]

Why methanol-free matters in practice: it removes a flammable-toxic inducer from the process, simplifies scale-up (no methanol feed strategy or off-gas handling), and eases regulatory filings where methanol classification is a burden. The literature shows PUPP/PDF reaching up to 9× the specific productivity of pGAP for CalB[12] — meaning "choosing the right methanol-free system" itself has large optimization headroom, exactly where the model's recommendation adds value. The platform's data already covers both methanol and methanol-free constructs (e.g. the learned pCS1 and pGAP effects), so it can directly answer "methanol-induced versus methanol-free constitutive, and which promoter," and give scenario-specific recommendations by product type, methanol-free requirement, and scale-up stage.

8. Product boundary

Evolrix AI's product boundary is recombinant proteins plus active molecules. The protein side spans industrial enzymes and therapeutic proteins; the active-molecule side extends the host into high-value small molecules via metabolic engineering. Concrete examples beyond the flagship case:

CategoryExamplesWhy Pichia
Industrial enzymesPhytase, lipase, cellulaseHigh-density secretion, disulfide folding
Therapeutic proteinsHSA, insulinEukaryotic PTM, GRAS-track host
Active small moleculesα-santalene (21.5 g/L), hyaluronic acid (0.8–1.7 g/L), 3-HPMevalonate / metabolic pathway engineering[6]

The active-molecule side is not aspirational: α-santalene has been produced at 21.5 g/L[6] in Pichia via an engineered mevalonate pathway, hyaluronic acid at 0.8–1.7 g/L, and 3-HP from methanol — showing the host's reach beyond proteins into high-value small molecules. Industrial enzymes like phytase, lipase, and cellulase are the bread-and-butter of Pichia manufacturing today; therapeutic proteins like human serum albumin (HSA) and insulin exploit the host's eukaryotic folding and secretion. The platform's scope is the union of these, not a single product class.

9. Commercialization & landing

Evolrix AI's commercial path targets three customer types, defined by who actually runs Pichia protein expression:

CustomerPain pointEntry
Research institutes — university/institute synthetic-biology and protein-engineering labsConstruct selection by trial-and-error; time-consuming literature searchLow-barrier tool + auditable evidence; build reputation and a data flywheel first
Pharmaceutical companies — biologics / recombinant-protein R&D teams (incl. CROs/CDMOs)Slow expression optimization; high screening costAccelerate expression-construct design, shortening the gene-to-first-product cycle
Synthetic-biology companies — industrial-enzyme / protein / biomanufacturing firmsRaising titer; lowering fermentation costConstruct/process recommendation + structure verification, owning titer and cost directly

The market these customers sit in is large. As above, the upstream synthetic-biology market is ~US$19.75B (2025) → ~US$56.48B (2031) at CAGR ≈19%[10]; the more direct protein-expression market is ~US$3.0–5.1B (2025) → ~US$5.0–7.7B (2031/2035) at CAGR ≈8–9%, with CROs/CDMOs the fastest-growing end-user and Asia-Pacific the fastest-growing region. A benchmark exists for the model: Cradle Bio serves exactly pharma (Novo Nordisk, J&J) plus chemicals/food/agriculture, confirming the reality of "research + pharma + synth-bio" customers; Evolrix AI's difference is a focus on the Pichia host and Chinese/patent data.

Referencing the validated Cradle Bio business model, a three-tier pricing structure is proposed:

TierFormFor
Entry / academicLow-cost or free web tool (evolrix.bio/science)Research institutes — reputation + data flywheel
Pro / APISubscription + full API into customer pipelines; customer data trains a private model, IP stays with the customerPharma / synth-bio R&D teams
Enterprise / projectPer-project / per-molecule + private deploymentLarge pharma / CDMO — accelerating R&D as a quantifiable value anchor

The quantifiable value propositions anchor these tiers: time saved (from "search literature + trial-and-error construct selection" to "model recommendation + evidence provenance"); higher titer / lower cost (construct/process recommendations own expression directly); trustworthy / auditable (every recommendation carries patent/literature provenance + structure verification, distinct from black-box tools); and a data moat → customer flywheel (customer wet-lab results feed back and continuously improve the model, on Evolrix AI's unique Chinese/patent data base). Specific pricing figures will be calibrated with early customers, as AI protein platforms are mostly enterprise/quote-based and Cradle/Basecamp do not publish unit prices.

10. Intellectual property

Evolrix AI is building an IP portfolio across software copyrights and invention patents, with filings in progress:

Two boundaries should be stated plainly alongside the commercial and IP claims — and we treat stating them plainly as a trust credential, not a limitation. A platform that can tell you precisely where its confidence ends is a platform you can build a regulated decision on. First, Boltz-2 structure outputs are predictions, not experimental structures — they are grounded and well-calibrated, but they are not crystallographic or cryo-EM data; the platform labels them as predictions and never presents them as measured. Second, expression-construct effects are directional evidence from cross-source correlation[5], not guaranteed outcomes — the platform distinguishes the computable axis (within-source ρ≈0.12, cross-source ρ≈0.26) from the wet-lab axis (absolute titer for a specific case), and flags the latter as requiring validation before any production claim. Claims are kept to what the evidence supports; predictions are labeled as predictions, and construct effects as directional evidence requiring wet-lab validation. In an industry where black-box "it will work" promises are the norm, this discipline is the basis on which a pharma or CDMO customer can actually adopt the output.

11. Roadmap

The full chain is sequence design → expression → fermentation → purification. The state of each layer is given as a capability map — what is validated today, what is in progress, what is planned — so that an investor or partner can see exactly where the platform's evidence sits and where the next inflection points are:

Design layer — validated

Manufacturing layer — planned, needs wet-lab partners

Near & mid term

The two layers are not separate pipelines; they are connected by an active-learning spine. As the manufacturing layer runs — real fermentation batches with measured titers and real purification yields — that outcome data flows back into the host-engineering dataset. A construct that over- or under-performs the platform's prediction becomes a new labeled record, sharpening the next ρ estimate and the next effect-size call. The design layer is validated today; the manufacturing layer is on the roadmap, connected to the same evidence-and-provenance spine so that the loop stays auditable end to end.

Scope. Platform whitepaper. Describes the Evolrix AI platform, its validated capabilities, its data sources, and its roadmap. Predictions are labeled as predictions; construct effects are labeled as directional evidence requiring wet-lab validation. FDA approvals cited here apply to Pichia-produced products (Kalbitor, Jetrea, Semglee), not to the Evolrix AI platform or the host itself. Generated 2026-07-04.
Run your own question through Kairos Research. Open Kairos →

References

  1. FDA Center for Drug Evaluation and Research. Biologics License Applications (BLA) and approvals: Kalbitor (BLA 125277), Jetrea (BLA 125422), Semglee (BLA 761201). U.S. Food and Drug Administration. FDA Biologics Approvals.
  2. Zha J, et al. Advances in Metabolic Engineering of Pichia pastoris Strains as Powerful Cell Factories. J Fungi (Basel), 2023. doi:10.3390/jof9101027.
  3. Market data: phytase market size (~$350M/yr) and Pichia commercialization scale (>5,000 proteins, >70 products, >300 licensed industrial processes, >300 biotech/pharma companies). Mordor Intelligence / MarketsandMarkets, 2025–2026 reports; Biomolecules 13(3):441 (2023); Front. Biosci.-Elite (2024).
  4. Evolrix AI Science. Structure verification live demo: ubiquitin wild-type pLDDT 0.93 vs scrambled 0.48 (Boltz-2 prediction). evolrix.bio/science.
  5. Evolrix AI host-engineering dataset (v0.5): ~627 regulatory records from 83 independent sources; 21 cross-source folds; within-source (de-confounded) Spearman ρ≈0.12; cross-source ρ≈0.26; overall ρ≈0.37; high-confidence folds xylanase/mannanase ρ 0.76–0.89; learned promoter pCS1 effect +0.7 (n=62); chaperones ERO1/SBH1 positive. evolrix.bio/research.
  6. α-Santalene production in Pichia pastoris at reported titer 21.5 g/L; hyaluronic acid 0.8–1.7 g/L; 3-HP from methanol. Zha J, et al. J Fungi 9(10):1027 (2023); Metabolic engineering of K. phaffii for 3-HP production from methanol, PubMed 39979934 (2025).
  7. Human lactoferrin expression data from 4 independent Chinese institutions: CN118147180A (ZenoBio), CN115960175A (Jiangnan University + Mengniu), CN116970503B (Jiangnan University), CN1718726A (Sanyuan). Chinese patent database (CNIPA / Google Patents).
  8. NCBI Bookshelf. Design-Build-Test-Learn: Impact of AI on the Synthetic Biology Process (The Age of AI in the Life Sciences) — DBTL canonical cycle; AI/ML entering the Design and Learn phases, foundries automating Build/Test.
  9. 15th Five-Year Plan (2026–2030) sources: CPC Central Committee Recommendations for the 15th Five-Year Plan (4th Plenum, 2025) + 15th Five-Year Plan outline (biomanufacturing as one of six prospective future industries, NDRC/Xinhua 2026-04); MIIT drafting the 15th Five-Year Biomanufacturing Development Plan (Xinhua 2025-12-19: flagship products + AI application cases, pilot-platform cultivation); 2025 Biomanufacturing Conference data (total ~¥1.1 trillion, >70% of global fermentation capacity); iCapital forecast 2030 China ~¥1.8 trillion, ~25% of global.
  10. Mordor Intelligence. Synthetic Biology Market 2026–2031 — US$19.75B (2025) → US$56.48B (2031), CAGR 19.14%; genome engineering 33.21% share, Asia-Pacific fastest-growing; AI integration named a core growth driver.
  11. Bi-directionalized promoter systems for methanol-free peroxygenases, Microb. Cell Fact. (2024) — methanol toxicity (2–5% v/v lethal to cells); Engineering the expression system for K. phaffii: a methanol-free expression system, PMC6736287 — GAP/DAS1 engineering, methanol-free routes.
  12. Bioprocess performance of novel methanol-independent promoters (PDF/PUPP), Microb. Cell Fact. (2021) — PUPP/PDF up to 9× PGAP for CalB; Enabling growth-decoupled K. phaffii production based on the methanol-free PDH promoter, PMC10076887.
  13. Jacobs P, et al. Engineering complex-type N-glycosylation in Pichia using GlycoSwitch. Nat. Protoc. 4:58–70 (2009); pichia.com GlycoSwitch — OCH1 knockout, humanized glycosylation, SuperMan5.
  14. Global synthetic-biology national strategies [R43]: US Executive Order 14081 (2022-09) + CHIPS & Science Act National Engineering Biology R&D Initiative + DoD BioMADE; UK National Vision for Engineering Biology (DSIT, 2023, £2B/10yr, one of five critical technologies); EU March-2024 biotech/biomanufacturing acceleration measures + EC bioeconomy country strategies (JRC, 2025-09); South Korea Synthetic Biology Promotion Act (passed 2025, effective 2026, world's first dedicated law, MSIT) + National Strategic Technology (2023-12); Japan Bioeconomy Strategy (¥100T market target). Economic scale: McKinsey The Bio Revolution and UK business.gov.uk engineering biology — $2–4T/yr global impact (2030–2040), ~$30T by end of century (~1/3 of global manufacturing). Reviews: SynBioBeta The Influence of Synthetic Biology on National Bioeconomy Strategies (2024), OECD synthetic-biology governance report.

Sources are listed in order of appearance. Patent data was mined from public patent databases (Google Patents, CNIPA). FDA BLA numbers are verifiable at the FDA biologics approvals page; FDA approvals apply to Pichia-produced products (Kalbitor, Jetrea, Semglee), not to the Evolrix AI platform or the host itself. Internal dataset statistics (pLDDT, ρ, effect sizes) are from the Evolrix AI host-engineering calibration dataset and are reproducible from the published records.