Benchmarks
Retrospective dry-lab design quality metrics — computed on the Kairos v0.4.0 snapshot and reported as of that snapshot. Both figures measure the scoring functions themselves (codon-reward correlation with real titers within a protein; element-effect correlation across sources), which have not been refit since. Candidate generation and ranking have changed since v0.4.0 — v0.7.0 made synthesis blockers outrank paper-optimality and v0.8.0 added synonymous-codon repair, both of which changed which candidate wins on the benchmark set. Every number is reproducible from the deterministic design tools — no opaque scoring.
66,000+ full-text papers · 627 records · 34 scored folds (21 cross-source) · 4 scored axes · claim-level audit
Host fidelity
Kairos scores each CDS with a host-fidelity reward (codon adaptation + genome-language-model host probability + host-mismatch guardrail). The ranking is checked against real data: for each of 5 proteins, the Spearman ρ between the reward and that protein’s own measured titers.
≈ +0.35
Spearman correlation between the Kairos host-fidelity reward and measured titers, computed within each protein and averaged across 5 proteins (28 real PichiaCLM variants). Design-quality ranking — not a titer prediction.
Per-protein Spearman
Reward-vs-titer correlation within each protein, averaged over 5 proteins. Retrospective on the PichiaCLM dataset; n=5 — per-protein values vary (one of five negative), so the mean is directional with a wide confidence interval.
Expression-construct optimization
Kairos learns element effects from expression-construct data. The learned effects are consistent with known biology: pCS1 and ERO1/SBH1 promoter/terminator elements show the expected directional impact. Cross-source generalization is measured across independent data cohorts.
≈ 0.057–0.265
Cross-protein construct ranking, re-pinned 2026-07-29: conservative within-source (deconfounded) ρ ≈ 0.057; cross-source ρ ≈ 0.265 (includes lab-scale effects); overall ≈ 0.328. LOGO evaluation.
~627 records
627 records; 34 scored folds, 21 cross-source (deconfounded mode: 22 scored, 9 cross-source). Retrospective calibrator evaluation; not a prospective predictor of expression or titer.
Coverage
Every design candidate is scored across multiple axes — each a deterministic, re-runnable computation:
Synthesis safety
Restriction sites, repeats, GC content, stop codons — hard constraints that block synthesis.
5′ translation initiation
Kozak context, secondary structure, start-codon accessibility.
Host codon fidelity
Codon adaptation to Pichia usage tables; within-protein reward-vs-titer ρ ≈ +0.35.
Developability
Protein-level flags: aggregation tendency, glycosylation sites, stability signals.
Evolution
Every reported number comes from a versioned, re-runnable snapshot. Eleven milestones from the public changelog — v0.1.0 to today.
Ask Kairos a research question — the dossier shows its evidence, scores, and provenance.