Minerva

Coevolutionary discovery with a genome language model

GitHub · Model card · Paper

DNA, FASTA, GenBank, protein, RNA, or Minerva's mixed format. Up to 4,096 tokens, about 10 kb of a bacterial locus.

Raw DNA and FASTA are gene-called with Pyrodigal and translated for you; GenBank CDS features are translated from the annotation; a protein is wrapped in a <+> strand marker; a mixed-format string (<+>MKV…<+>acgt…<->…) is used as is. One token is one amino acid, one nucleotide, or one strand marker, so a typical bacterial locus (~88 % coding) fits about 10 kb in 4,096 tokens and pure non-coding DNA about 4 kb. Longer input is cut at a gene boundary and the summary says what was kept. Heads: RNA base pairing, DNA repeats and protein contacts.

Examples

Output

Citation: Coevolutionary mining of prokaryotic non-coding elements with a genome language model Li, D.B. & Brixi, G. et al. bioRxiv (2026); doi: 10.64898/2026.09.22.753630