Benchmarks

Benchmarks & methods

Retrospective dry-lab design quality metrics — computed on the Kairos v0.4.0 snapshot and reported as of that snapshot. Both figures measure the scoring functions themselves (codon-reward correlation with real titers within a protein; element-effect correlation across sources), which have not been refit since. Candidate generation and ranking have changed since v0.4.0 — v0.7.0 made synthesis blockers outrank paper-optimality and v0.8.0 added synonymous-codon repair, both of which changed which candidate wins on the benchmark set. Every number is reproducible from the deterministic design tools — no opaque scoring.

66,000+ full-text papers · 627 records · 34 scored folds (21 cross-source) · 4 scored axes · claim-level audit

Important: These are dry-lab design quality ranking metrics. They measure how well Kairos ranks construct designs relative to each other. They are NOT expression level predictions, titer predictions, or wet-lab outcome guarantees. Real expression requires experimental validation.

Host fidelity

Within-protein host fidelity

Kairos scores each CDS with a host-fidelity reward (codon adaptation + genome-language-model host probability + host-mismatch guardrail). The ranking is checked against real data: for each of 5 proteins, the Spearman ρ between the reward and that protein’s own measured titers.

Within-protein ρ vs real titers

≈ +0.35

Spearman correlation between the Kairos host-fidelity reward and measured titers, computed within each protein and averaged across 5 proteins (28 real PichiaCLM variants). Design-quality ranking — not a titer prediction.

Method

Per-protein Spearman

Reward-vs-titer correlation within each protein, averaged over 5 proteins. Retrospective on the PichiaCLM dataset; n=5 — per-protein values vary (one of five negative), so the mean is directional with a wide confidence interval.

Expression-construct optimization

Learned element effects

Kairos learns element effects from expression-construct data. The learned effects are consistent with known biology: pCS1 and ERO1/SBH1 promoter/terminator elements show the expected directional impact. Cross-source generalization is measured across independent data cohorts.

Cross-source ρ

≈ 0.057–0.265

Cross-protein construct ranking, re-pinned 2026-07-29: conservative within-source (deconfounded) ρ ≈ 0.057; cross-source ρ ≈ 0.265 (includes lab-scale effects); overall ≈ 0.328. LOGO evaluation.

Dataset

~627 records

627 records; 34 scored folds, 21 cross-source (deconfounded mode: 22 scored, 9 cross-source). Retrospective calibrator evaluation; not a prospective predictor of expression or titer.

Fig. 1 — Spearman ρ between Kairos scores and real expression data. Within-protein host fidelity +0.35 (directional, wide CI). Construct-ranking figures re-pinned 2026-07-29. Not a titer prediction.

Coverage

What Kairos evaluates today

Every design candidate is scored across multiple axes — each a deterministic, re-runnable computation:

Axis 1

Synthesis safety

Restriction sites, repeats, GC content, stop codons — hard constraints that block synthesis.

Axis 2

5′ translation initiation

Kozak context, secondary structure, start-codon accessibility.

Axis 3

Host codon fidelity

Codon adaptation to Pichia usage tables; within-protein reward-vs-titer ρ ≈ +0.35.

Axis 4

Developability

Protein-level flags: aggregation tendency, glycosylation sites, stability signals.

Kairos version: v0.19.6 · Engine: v0.1.0
Metrics snapshot: v0.4.0 (scoring functions not refit since; generation/ranking changed in v0.7.0 & v0.8.0)
Data snapshot: 2026-07-29 (re-pinned)
Method: retrospective LOGO cross-validation on existing expression-construct data
Scope: Pichia (Komagataella phaffii) only
Full changelog →
Fig. 2 — Retrospective LOGO cross-validation on existing expression-construct data. Every reported ρ is computed on held-out sources.

Evolution

One snapshot, on a trajectory

Every reported number comes from a versioned, re-runnable snapshot. Eleven milestones from the public changelog — v0.1.0 to today.

v0.1.02026-07-01explainable Pichia CDS design
v0.3.02026-07-01chat UX + design-panel from a name
v0.4.02026-07-09auditable research workbench
v0.5.02026-07-24review layer + artifact live
v0.8.02026-07-26synonymous-codon repair — build the winner
v0.12.02026-07-28audit checks citation support
v0.15.02026-07-29email sign-in live
v0.17.02026-08-14manufacturability panel
v0.18.02026-08-21three-layer conversation
v0.18.52026-08-28chat summary + provenance manifest
v0.19.02026-08-29hypothesis-native verification

Every number here is re-runnable. Try one

Ask Kairos a research question — the dossier shows its evidence, scores, and provenance.