Inspiration

We were inspired by DNA visualizers available online. We noticed they display a lot of information in a format that is hard for non-experts to read and understand. Our goal was to create a DNA visualizer that would be suitable for anyone to get an at-a-glance idea of what real human DNA looks like.

What it does

Our project takes the human genome reference from the National Library of Medicine and organizes it by chromosome. 22 autosomes, the X chromosome, and the Y chromosome are indexed and stored to be displayed and manipulated as necessary. You can navigate through a visual representation of each of these chromosomes and see the base pairs at every location. You can search for genes by name and find the chromosome and nucleotides at which they live.

How we built it

BioPython indexes and sequences the human reference genome (GRCh38). This turns 3.1 billion base pairs into something viable. A PostgreSQL database organizes an annotated genome file so we can quickly look up genes and their locations in a specific chromosome. Ursina is the graphics library we used to build a UI to interact with our genome and annotated genome information.

Challenges we ran into

The human reference genome is more than 3 GB of pure text. The annotation file that accompanies it is another 2 GB of text. We ran into challenges concerning memory limitations and timely processing of information. Trying to store the reference genome in more complex data structures was estimated to take up to more than 250 GB of RAM in some scenarios. With our first solution, the window would be unresponsive for 2 minutes as the genome was sequenced.

Accomplishments that we're proud of

We were able to handle multiple large sets of information relatively quickly and with good memory efficiency. We implemented a database to be able to rapidly search 3.1 billion base pairs for locations of specific genes.

What we learned

As computer science majors, we knew very little about how DNA was organized and stored, so we had to educate ourselves on the basics of DNA organization. We researched libraries to manipulate the genomes we found, and we found accurate genomes from reputable government organizations.

What's next for DNA Visualizer

A goal that we were unfortunately not able to implement within the allotted time was a mutation lookup function. Taking data on mutations and their locations, we would have been able to scan for and identify healthy copies of genes versus mutated copies of genes.

Built With

Share this project:

Updates

Submission history