There is a temptation, when you build an AI design tool, to claim it predicts expression. The claim sells. It is also, today, not honestly defensible. Expression in Pichia depends on folding, glycosylation, secretion efficiency, strain background, copy number, promoter strength, methanol feeding, temperature, and a dozen other variables no in-silico tool can fully account for. A design engine that claims to predict all of that is overpromising.
Kairos does something narrower and, we believe, more useful: it ranks dry-lab design quality — the properties of the coding sequence and the construct that can be computed deterministically.
What Kairos scores
Four axes, each computed from the printed sequence:
- Synthesis safety — will the CDS survive gene synthesis without trouble? (GC, homopolymers, repeats, restriction sites, forbidden motifs.)
- Host resemblance — does the CDS look like something K. phaffii would actually write? (Codon usage + %MinMax rhythm.)
- 5′ translation initiation — is the start region accessible to the ribosome?
- Protein developability — what are the intrinsic liabilities of the target protein itself? (Length, MW, pI, cysteines, glycosylation sequons, hydrophobicity.)
What Kairos does not predict
Expression level. Titer. Secretion efficiency. These depend on variables outside the sequence — strain, process, conditions — and require wet-lab measurement. Kairos will not give you a number for them, because producing such a number would be persuasion dressed as prediction.
A reserved slot for the truth
Instead of pretending the gap doesn't exist, the design provenance reserves a feedback slot for wet-lab assay data. When expression data comes back from the bench, it can be attached to the design record — closing the loop from dry-lab ranking to measured outcome, honestly. This is the experiment-ready design of the system: not a claim, but a place for the data to land.
Why this matters
A tool that overstates its scope erodes trust the first time a “high-scoring” sequence expresses poorly. A tool that states its scope honestly earns trust every time a well-designed sequence does express — because the user understands what the score meant, and what it didn't. Provenance, not persuasion, starts with saying what you don't know.
References
- Sharp PM, Li WH. The codon Adaptation Index — a measure of directional synonymous codon usage bias. Nucleic Acids Res, 1987. PMID:3547335
- Gustafsson C, et al. Codon bias and heterologous protein expression. Trends Biotechnol, 2004. PMID:15245907
- Zha J, et al. Advances in Metabolic Engineering of Pichia pastoris Strains as Powerful Cell Factories. J Fungi (Basel), 2023. PMID:37888283
- Maity N, et al. Statistically Designed Medium Reveals Interactions between Metabolism and Genetic Information Processing for Production of Stable Human Serum Albumin in Pichia pastoris. Biomolecules, 2019. PMID:31590267