Research library
Models.
The models behind our research products and design programs.
Browse by scientific role. Each page says what the model is for, where it fits, and what its outputs mean.
01
Foundation models
cdsBERT
A codon-aware protein modeling research direction for understanding coding-sequence effects.
Dual Triangle Attention
A foundation-model research direction for stronger position-aware bidirectional sequence modeling.
ESMC
EvolutionaryScale Biohub's ESM Cambrian model family, exposed through FastPLMs as ESM++ checkpoints.
Vec2Vec
Aligns protein, annotation, and language-model representations so biological knowledge can move between spaces.
04
Oracle models
EC General
Suggests broad enzyme-class hypotheses from protein sequence.
EC Rigor
Provides stricter enzyme-function hypotheses when higher-confidence EC assignment is needed.
E. coli Expression
Prioritizes proteins that are more likely to express successfully in E. coli workflows.
Homodimer
Estimates whether a protein sequence is likely to self-associate as a homodimer.
Atlas Oracle Suite
A set of fast protein property predictors for triaging sequence quality, function, localization, and developability.
kcat
Estimates enzyme turnover potential from protein sequence.
pH Preference
Estimates the pH range where a protein may be most compatible or active.
Realness
Scores whether a sequence looks like a natural protein rather than random or poorly formed sequence.
Solubility
Prioritizes protein sequences that are more likely to remain soluble.
SoluProt
Provides a complementary solubility-style readout for protein developability triage.
Subcellular Localization
Predicts likely cellular localization signals from protein sequence.
Taxon
Estimates taxonomic signal in a protein sequence for context, quality control, and dataset review.
Temperature Stability
Estimates whether a protein sequence is likely to tolerate higher-temperature conditions.
Translator
Turns protein sequences into structured functional annotation hypotheses.