We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

We were inspired by our strengths and ambitions to produce the best project we could in this challenge. Friday night seminar by Bryan and Mark discussed the challenges of using large .bed data representations of genomic sequences. They had three challenges, one for comparing two sequences, one for adding or subtracting overlaps to understand the sensitivity to the sequence positions. The last challenge was to use multiple genomic experiments to understand the statistical variability of the processes. This requires multiple input files and sampling of the overlap regions the (n,k) problem. The computational effort for multiple genomic experiments was presented as challenge, perhaps taking hours to run this matching problem. Our code runs in 17 seconds.

What it does:

We input .bed files, compute the genomic sequence overlap, and output a .bed file. The multiple file sequencing algorithm is fast and solves the (n,k) method as required. We solved the first tow input challenge, the second overlap challenge, and the multiple three input challenge. Our program runs in parallel and would work well in server or cloud environments.

How we built it

Nick and Phil worked on programs to analyze the data and produce .bed and .json files for visualizing. Nick wrote a script in Python 3 and Phil wrote an algorithm in Fortran and MPI parallel. Mike created a Flask project to visualize our results on the web in matplotlib plots. Victoria worked on our video, its script, and an attempt at visualizing our data with D3.JS.

Challenges we ran into

D3.js proved very challenging with a steep learning curve we weren't able scale it. The team developed a good working relationship.

Accomplishments that we're proud of: The fast and powerful genomic comparison code will expand the user community ability to run multiple validation experiments and get good analysis results. At least 10x over existing methods in use today.

What we learned

Nick learned sleep is for the weak as he stayed up all night working on his amazing Python code which will never ever be as fast as Phil's Fortran algorithms. Nick also learned more about creating custom classes in Python.

What's next for DIG: Doing Important Genomics- HATCh #1, expand the visual ability and move the application to the cloud for general HA use.

Built With

Share this project:

Updates

Submission history