Inspiration

  • The Problem: Researchers spend hours manually cross-referencing NCBI records, calculating base compositions, translating reading frames, and cleaning data for downstream pipelines.
  • The Insight: Generative AI alone can hallucinate biological metrics, while traditional tools lack contextual reasoning. The solution lies in pairing deterministic bioinformatics algorithms with guardrailed generative AI.
  • The Result: A serverless, sub-second bioinformatics engine that transforms any NCBI accession ID into verified biochemical facts, downloadable FASTA files, and structured biological breakdowns.

What It Does

  • Instant NCBI Resolution: Ingests any accession ID and dynamically fetches complete FASTA sequence records and NCBI entry definitions on the fly.
  • Deterministic Bioinformatics Algorithms: Automatically calculates GC/AT percentages, 3-frame translated ORFs, residue frequencies, and peptide-condensation-corrected molecular weights.
  • 1-Click Sequence Export: Provides direct streaming downloads for raw sequence .fasta files and translated polypeptide open reading frames.
  • Guardrailed GenAI Insights: Uses Amazon Bedrock (Nova Lite) to generate structured biological explanations, protected by AWS Bedrock Guardrails for enterprise safety and PII filtering.
  • Sub-Second Serverless Delivery: Operates as a serverless AWS application (API Gateway v2 + Lambda) delivering both an interactive web dashboard and a headless JSON REST API.

How We Built It

We designed and built the platform from the ground up on AWS using modern cloud-native, serverless, and generative AI technologies.

Architecture Block Diagram

       ┌───────────────────────────────┐
       │ User / Web Browser / REST API │
       └───────────────┬───────────────┘
                       │
                       │ (HTTP Request)
                       ▼
         ┌───────────────────────────┐
         │  Amazon API Gateway v2    │
         └─────────────┬─────────────┘
                       │
                       │ (Low-Latency Proxy)
                       ▼
     ┌───────────────────────────────────┐
     │ AWS Lambda Engine (Python 3.11)   │
     └───────────┬───────────────────┬───┘
                 │                   │
  (Bi-directional)│                   │ (Invokes Agent / Direct Tooling)
                 ▼                   ▼
   ┌──────────────────────────┐┌────────────────────────────────────────┐
   │ NCBI Entrez E-Utilities  ││ Strands Agent Runtime (agent.py)       │
   │  (efetch / esearch)      ││  ├─ Bedrock AgentCore App Runtime      │
   └──────────────────────────┘│  ├─ Strands Bio Agent (Tool Routing)   │
                               │  └─ Modular Tool: analyze_sequence     │
                               └───────────────────┬────────────────────┘
                                                   │
                                                   │ (Guardrailed Reasoning)
                                                   ▼
                               ┌────────────────────────────────────────┐
                               │ Amazon Bedrock (Guardrail + Nova Lite) │
                               └────────────────────────────────────────┘

1. Cloud Infrastructure as Code (IaC)

  • AWS CDK (Cloud Development Kit): Scripted the entire architecture in Python using aws_cdk, deploying the CloudFormation stack testbio-agent-app into us-west-2.
  • Amazon API Gateway v2 (HTTP API): Configured routes (GET /, GET /download-fasta, and GET and POST /sequence-analysis) with automatic CORS handling and low-latency payload proxying.

2. Serverless Bioinformatics Engine (AWS Lambda)

  • Real-time NCBI E-Utilities: Implemented an HTTPS client integrating NCBI efetch.fcgi and esearch.fcgi to stream FASTA sequence records and entry definitions directly from NCBI's nuccore and protein repositories on demand.
  • Deterministic Calculations: Built a high-performance computation engine that calculates GC/AT percentages, performs 3 forward reading frames codon scanning to isolate Open Reading Frames (≥ 30 aa), and computes peptide-condensation-corrected molecular weights: $$\text{MW} = \sum \text{AA}_i - (N - 1) \times 18.015 \text{ Da}$$
  • 1-Click Streaming FASTA Exporter: Engineered in-memory file generation emitting standard Content-Disposition: attachment headers for instant sequence downloads without needing S3 buckets.

3. Generative AI and Safety Guardrails (Amazon Bedrock)

  • Amazon Nova Lite (amazon.nova-lite-v1:0): Orchestrated structured biological reasoning via Bedrock's Converse API, converting raw metrics into peer-review-quality insights.
  • Active Bedrock Guardrail (gtbuv74mobh6): Applied enterprise safety filters enforcing PII masking (SSNs, cards, credentials) and prompt injection safeguards.
  • Strands Agent Runtime (agent.py): Integrated the Strands agent framework (strands-agents) wrapped with AWS Bedrock AgentCore (BedrockAgentCoreApp), registering deterministic bioinformatics tooling (@tool(analyze_sequence)) for modular, autonomous genomic reasoning.

Challenges We Ran Into

1. LLM Output Truncation and Token Headroom

  • The Challenge: During early tests on complex accession IDs (such as the plant ribosomal cluster MN908990), the model’s concluding paragraph was cut off mid-sentence ("In summary, this sequence from Ph...") due to a tight token ceiling (maxTokens: 600).
  • The Solution: Diagnosed the stopReason: max_tokens error in the Bedrock response, tested token usage locally across various sequences, and removed the artificial maxTokens cap to enable model-default unlimited token generation (stopReason: end_turn), allowing the full biological summary to conclude naturally.

2. Preventing AI Hallucinations via Deterministic Anchoring

  • The Challenge: Large language models frequently hallucinate or approximate mathematical calculations like exact nucleotide base counts, codon frame translations, and molecular weights.
  • The Solution: Implemented a hybrid architecture. The deterministic Python engine computes all biological math first. These verified facts are then passed into Bedrock's prompt as immutable ground truth, allowing the LLM to focus purely on biological reasoning and context rather than arithmetic.

3. Multi-Database Accession Disambiguation

  • The Challenge: NCBI accession IDs do not always explicitly declare whether they represent a nucleotide sequence (DNA/RNA) or a translated protein (polypeptide).
  • The Solution: Built an automated database resolution waterfall that attempts queries across both nuccore and protein databases, detects sequence alphabets, and dynamically routes the data to the appropriate nucleotide or protein analysis pipeline.

4. Direct FASTA Downloads in Serverless APIs

  • The Challenge: Enabling immediate, one-click file downloads for researchers without introducing the latency, storage costs, and permission complexities of temporary S3 presigned URLs.
  • The Solution: Engineered on-the-fly streaming directly through API Gateway and Lambda using HTTP attachment headers (Content-Disposition: attachment; filename="<accession>.fasta"), delivering zero-storage, instant file downloads.

5. Guardrail Calibration for Scientific Nomenclature

  • The Challenge: Strict enterprise PII and content moderation filters can occasionally flag complex biological nomenclatures, accession hashes, or genetic codes as sensitive data.
  • The Solution: Structured the system and user message boundaries in the Bedrock Converse API payload to ensure guardrails cleanly inspect input and output text without triggering false positives on scientific terminology.

6. Massive Whole-Chromosome Streaming and Serverless Memory Bounds

  • The Challenge: Querying whole-chromosome reference sequences (such as Human Chromosome 11 NC_000011, spanning over 135 million base pairs) can easily overwhelm serverless memory limits, saturate network sockets, and cause API Gateway 29-second timeout failures when performing unbounded reads.
  • The Solution: Engineered a bounded stream-reading pipeline (MAX_STREAM_BYTES = 5 MB) using chunked HTTP streams. This retrieves the complete entry metadata and millions of base pairs within seconds, ensuring rapid, crash-proof analysis even for massive eukaryotic genomes.

7. Dynamic Header Normalization and Version Disambiguation

  • The Challenge: NCBI and UniProt FASTA headers frequently return complex, versioned identifiers (e.g., querying NR_046261 returns >NR_046261.1 Sus scrofa...) or pipe-delimited prefixes (>sp|P04637.4|P53_HUMAN...). Naive header slicing caused accession versions and database prefixes to leak into definitions and UI display descriptions (e.g., displaying "NR_046261.1 and so on").
  • The Solution: Developed a generic, zero-hardcoding header parser that dynamically isolates the user-queried accession, separates version suffixes, extracts clean definition sentences, and formats canonical FASTA export headers cleanly across any nucleotide or protein record.

Accomplishments That We're Proud Of

  • Built a Zero-Hallucination Hybrid Architecture: We successfully bridged deterministic bioinformatics algorithms with generative AI. By ensuring that all base counts, GC/AT ratios, ORF translations, and molecular weights are calculated deterministically first, we eliminated the mathematical hallucinations common in pure LLM applications.
  • Full Production Serverless Deployment on AWS: Deployed a scalable, cost-efficient cloud architecture using AWS CDK in us-west-2, combining Amazon API Gateway v2, Python 3.11, AWS Lambda, and Amazon Bedrock with zero idle infrastructure costs.
  • Instant FASTA Streaming: Engineered an in-memory, zero-storage file streaming pipeline that generates and downloads standard FASTA formatted assets instantly for the user.
  • Resilient Across Extreme Sequence Scales: Successfully validated and benchmarked queries spanning from short genomic markers (400 bp) and viral polyproteins (SARS-CoV-2 YP_009724389) up to massive multi-megabase mammalian chromosomes (NC_000011).

What We Learned

  • Ground-Truth Anchoring is Essential for Scientific AI: Large language models excel at synthesizing biological function and clinical context, but should never be tasked with raw arithmetic or exact substring indexing. Passing pre-computed deterministic facts as immutable prompt context produces rock-solid, peer-review-quality AI summaries.
  • Serverless Optimization for Bioinformatics Data: Streaming payloads directly through API Gateway v2 and AWS Lambda with in-memory byte buffers avoids the latency and permissions friction of intermediary storage (like S3), keeping response times sub-second for standard queries.
  • Fine-Tuning Enterprise Guardrails: Calibrating Amazon Bedrock Guardrails requires careful message separation so enterprise filters (like PII masking and prompt injection shields) inspect conversational context without flagging biological IUPAC sequences or accession codes as false positives.

What's Next for the Platform

  • Pairwise and Multiple Sequence Alignment (BLAST / Clustal Omega): Enable researchers to compare two or more accession IDs side-by-side with automated visual alignment and variant calling.
  • 3D Protein Structure Visualization: Integrate Mol* or py3Dmol directly into the web dashboard to fetch and render interactive 3D structures from AlphaFold and PDB for protein accessions.
  • Full 6-Frame Translation and Batch Export Pipeline: Expand ORF analysis to include reverse-complement frames (frames -1, -2, -3) with codon optimization scoring for recombinant protein expression, paired with automated multi-sequence ZIP archive generation.

Built With

  • amazon-bedrock
  • amazon-nova-lite
  • aws-bedrock-guardrails
  • aws-cdk
  • aws-lambda
  • amazon-api-gateway-v2
  • python-3.11
  • ncbi-entrez-e-utilities
  • strands-agents
  • bioinformatics
  • html5-css3

Built With

  • amazon-api-gateway-v2
  • amazon-bedrock
  • amazon-nova-lite
  • aws-bedrock-guardrails
  • aws-cdk
  • aws-lambda
  • bioinformatics
  • ncbi-entrez-e-utilities
  • python
  • strands-agents
Share this project:

Updates

Submission history