Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Well, this is actually the old way of parsing a genome. What you've described is essentially a Hidden Markov Model with discrete states, which is the bread and butter of genome sequencing.


I'm actually curious about that distribution; could you gain compression efficiency by grouping them into 3-base-pair codons? Or is DNA pretty much random at the base-pair level and the codon redundancy makes it actually work anyway?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: