MinervaCoevolutionary discovery with a genome language model
GitHub · Model card · Paper
DNA, FASTA, GenBank, protein, RNA, or Minerva's mixed format. Up to 4,096 tokens, about 10 kb of a bacterial locus.
Raw DNA and FASTA are gene-called with Pyrodigal and translated for you; GenBank CDS features are
translated from the annotation; a protein is wrapped in a <+> strand marker; a mixed-format string
(<+>MKV…<+>acgt…<->…) is used as is. One token is one amino acid, one nucleotide, or one strand
marker, so a typical bacterial locus (~88 % coding) fits about 10 kb in 4,096 tokens and pure
non-coding DNA about 4 kb. Longer input is cut at a gene boundary and the summary says what was
kept. Heads: RNA base pairing, DNA repeats and protein contacts.
Output
Citation: Coevolutionary mining of prokaryotic non-coding elements with a genome language model Li, D.B., Brixi, G., Kim, A.S., Fiamenghi, M.B., Driscoll, C.L., Evans, S.A., Gao, A., Ivanova, N.N., Kyrpides, N.C., Deisseroth, K., Wilkinson, M.E., Fischbach, M.A. & Hie, B.L. bioRxiv (2026); doi: 10.64898/2026.09.22.753630