Inspiration

What it does# ExperimentMind

The Problem

Data scientists lose the majority of their working hours to experiment bookkeeping — not science. Every training run follows the same manual cycle: write a config, launch the job, watch logs, copy metrics into a spreadsheet, compare against prior runs, write a Slack summary. None of this requires a human brain. Yet no tool has autonomously owned the full experiment lifecycle end-to-end.

MLflow and Weights & Biases are excellent trackers — but they are passive ledgers. They record what you tell them to record. You still have to do everything else manually.

ExperimentMind acts.


What It Does

Drop a YAML experiment config into a watched folder. Walk away. Wake up to a ranked leaderboard and a natural-language digest in your inbox. The agent handles everything in between.

The only times ExperimentMind surfaces to you are when a genuine decision is needed:

  • A new best model is ready for deployment approval
  • A config has an error that requires human judgment
  • A job fails repeatedly and needs investigation

Everything else runs silently in the background.


How We Built It

ExperimentMind is a five-agent pipeline built with the AWS Strands Agents SDK, connected to Amazon Bedrock Mantle (GLM 4.7 Flash via AWS credits), with a FastAPI backend and a real-time mission control dashboard.

The Five Agents

Config Watcher Agent monitors experiments/queue/ using the watchdog library. When a new YAML file is detected, it immediately triggers the validation pipeline. Tools: detect_new_config, list_queue_configs.

Validator Agent checks the config against a schema — required fields, valid metric names, accessible dataset paths — and routes the config to running/ on success or failed/ with a detailed error report on failure. Tools: read_yaml_config, validate_required_fields, check_paths, move_config_to_running, move_config_to_failed.

Job Launcher Agent executes the training job as a subprocess, monitors the running process by tailing logs, detects completion or failure, and identifies failure reasons (OOM, bad gradient, data errors). Tools: run_local_job, wait_for_completion, tail_logs.

Results Logger Agent parses training output — extracting metrics from JSON or plaintext log formats — and writes everything to a SQLite knowledge base via SQLAlchemy. Tools: parse_metrics_from_log, write_run_to_db, move_config_to_done, query_recent_runs.

Analyzer + Reporter Agent queries the full experiment history, ranks all runs by the primary metric, detects performance regressions across the last three runs, generates a natural-language interpretation using Bedrock Mantle, and delivers a formatted daily digest via email, Slack, or file. Tools: query_all_runs, compute_leaderboard, detect_regression, get_experiment_summary, format_digest, send_email_digest, save_digest.

Key Architectural Decision

Early in the build, small LLMs proved unreliable at multi-tool orchestration — given a sequence of tools to call, they would sometimes skip steps or call them out of order.

The solution: make every pipeline step deterministic Python, and reserve LLM reasoning only for the one task where it genuinely adds value — generating the natural-language interpretation of what experiment results mean.

The validation pipeline doesn't ask an LLM to decide what to do. It calls tools directly in sequence with explicit conditional routing. The LLM is invoked exactly once per pipeline run: to turn a ranked leaderboard and trend data into a readable paragraph that a data scientist can act on.

This produced a system that is both reliable and intelligent.

Amazon Bedrock Mantle Integration

ExperimentMind connects to models via Amazon Bedrock Mantle — AWS's OpenAI-compatible inference endpoint. Using Strands' OpenAIModel provider pointed at the bedrock-mantle regional endpoint, the system runs entirely on AWS credits with no external API dependencies.

Switching from local development (Ollama + Qwen 2.5) to production (Bedrock Mantle + GLM 4.7 Flash) required changing exactly three lines in agents/base.py. That's the composability Strands enables.


Challenges

Model orchestration reliability — Small LLMs cannot reliably sequence tool calls. Solving this required the architectural insight of separating deterministic pipeline execution from LLM reasoning entirely.

AWS quota management — New AWS accounts face daily token limits on Bedrock. Discovering Amazon Bedrock Mantle's OpenAI-compatible API was the breakthrough that resolved this, providing access to a broader model catalog without the quota constraints of the standard Bedrock ConverseStream API.

Pipeline robustness — Building an autonomous system that runs overnight without human oversight required careful error handling at every stage: config validation with detailed error reports, job failure detection with retry logic, and a non-blocking attention queue that surfaces issues without halting the pipeline.


What We Learned

Reserve LLM reasoning for judgment, not execution. The most reliable agentic systems use LLMs for what they're genuinely good at — interpreting data and generating language — and use plain Python for sequencing, routing, and deterministic operations.

Tool granularity matters. One tool, one thing, one structured return value. When tools do multiple things, debugging becomes impossible. The @tool decorator in Strands makes atomic tools easy to implement.

Abstract the model provider from day one. A three-line get_model() factory function saved hours of refactoring and made the switch from local to cloud inference seamless.


What's Next

  • SageMaker and Kaggle integration for cloud-scale training jobs
  • Automated hyperparameter suggestion based on experiment history
  • Multi-user support with team experiment sharing
  • Integration with existing MLflow and Weights & Biases tracking

Built With

AWS Strands Agents SDK · Amazon Bedrock Mantle · GLM 4.7 Flash · FastAPI · SQLite · SQLAlchemy · Watchdog · Python 3.12

How we built it

Challenges we ran into

Accomplishments that we're proud of

What we learned

What's next for ExperimentMind

Built With

Share this project:

Updates

Submission history