Codon optimization has long been treated as a single-answer problem: feed in a protein, get back one “optimized” sequence. Kairos rejects that frame. A coding sequence is a design decision with trade-offs — between manufacturability, host resemblance, translation initiation, and protein-level liabilities — and a single answer hides those trade-offs behind a score you cannot inspect.
Kairos is an AI research copilot for synthetic biology (Pichia-scoped, v0.3). It accepts natural-language questions, retrieves cited evidence, reasons with an LLM, and returns a panel of verifiable CDS candidates — each provably encoding the target protein, scored with host-fidelity signals, and accompanied by 3D structure predictions.
A panel, not a single answer
For each design run, Kairos proposes multiple synonymous coding sequences using a mix of transparent host heuristics and a self-developed, genome-learned model (EvoCodon). The candidates compete on the same scorecard. There is no auto-winning entry — the ranking reflects the data, and you see why each candidate lands where it does.
Four scoring axes
Every candidate is scored on the same four axes, each fully explainable:
- Manufacturability — GC window, homopolymers, tandem repeats, restriction sites, and forbidden synthesis motifs. A candidate either passes or carries explicit flags.
- Host-likeness — a genome-derived K. phaffii codon model that combines codon usage with %MinMax rhythm, deliberately rejecting naive max-CAI over-optimization.
- 5′ translation initiation — start-proximal mRNA accessibility. An open start region signals better ribosome recruitment.
- Protein developability — length, molecular weight, pI, cysteine count, N-glycosylation sequons, and hydrophobicity — the intrinsic liabilities of the target itself.
Multi-objective ranking
The panel is ranked by Pareto dominance across all four axes — never collapsed into one fragile aggregate score. The recommended candidate sits in the top Pareto tier and is clean of hard manufacturability violations. When a flag appears on the winner (for example, an extreme pI), it is an intrinsic property of the target protein, not a design defect — and Kairos says so plainly.
Correct by construction
Every candidate is verified by translation: it must encode the exact input protein, in-frame, with no premature stop. The design step calls verifiable tools to produce sequences — it never invents them. This is what we mean by provenance, not persuasion: every score is a deterministic computation you can re-run over the printed sequence.
References
- Sharp PM, Li WH. The codon Adaptation Index — a measure of directional synonymous codon usage bias, and its potential applications. Nucleic Acids Res, 1987. PMID:3547335
- Gustafsson C, et al. Codon bias and heterologous protein expression. Trends Biotechnol, 2004. PMID:15245907
- Ahmad M, et al. Efficient Expression of Lactone Hydrolase Cr2zen for Scalable Zearalenone Degradation in Pichia pastoris. Toxins (Basel), 2025. PMID:41591157
- Zha J, et al. Advances in Metabolic Engineering of Pichia pastoris Strains as Powerful Cell Factories. J Fungi (Basel), 2023. PMID:37888283