The DataDeck Agent exists to remove the run-up to analysis. Point it at a raw export and it hands back data that isn't just tidier but ready — consistent, validated, and already characterized — so you begin at the starting line of the actual work instead of clearing the track first. By the time it finishes, you know the shape of the data: what each field holds, how values are distributed, which relationships stand out, and where the soft spots are.
It is reachable through a CLI, an HTTP API, an SDK, an MCP server, and a chat bot, and it's deployable as a scale-to-zero cloud service. Built on the AWS Strands Agents SDK (TypeScript)
How it gets you there
It cleans by meaning, not by surface pattern. The agent works out what each column actually represents and reasons a per-column plan: it collapses values that refer to the same thing (USA / U.S.A. / United States → one), settles ambiguous date order, normalizes numbers and currency (converting across currencies only when supplied with rates), and sets the validation rules a column should hold to. A model decides the intent; a deterministic engine applies that plan identically on every run. That two-layer split is the point — the reasoning that rules can't do, over an execution core that's repeatable and needs no credentials.
It reaches that ready state as a series of gated decisions, not a bulk overwrite. Every value is routed one of three ways — fix, flag, or leave — and the changes that carry risk pass through approval gates. Ambiguous dates, mixed currencies, probable outliers, and gaps in otherwise-complete columns are held with a short reason and take effect only once cleared. So the data you sit down to analyze is settled, not assumed.
What each run produces
A complete, self-documenting deliverable set:
- The cleaned dataset (
cleaned.csv) — analysis-ready, with fixes applied and duplicates merged. - A cleaning report (
cleaning-report.html) — every change that was made, every value held for review, each with its reason, plus a column-by-column profile (types, fill rate, uniqueness, formatting issues). It reads as an account of exactly what happened to your data and why. - A pre-analysis report (
analysis-report.html) — your first orientation in the data: summary statistics, per-column distributions, and ranked correlations, with the highlights called out. - Machine-readable summaries (
summary.md,summary.json) — the same facts for humans and for downstream tooling. - All of the above packaged together (
<name>-datadeck.zip) so a run is one shareable artifact.
Who it's for
- Analysts & data scientists who want the pre-analysis grind gone and to open a dataset already oriented.
- Data & platform engineers who want a repeatable readiness step in a pipeline
- Developers & AI builders who want to call the capability (API/SDK), wire it into an editor/assistant (MCP), or run it as a hosted service
- Ops, finance, and support teams who live in exported spreadsheets and need them made analysis-ready fast, with a clear record of what was changed and what to check.
How it fits together
A Strands Agent with five tools (profile_dataset → inspect_columns → apply_cleaning_plan, plus quick_auto_clean and list_sheets) drives a pure core engine (parse → detect → clean/applyPlan → analyze → report → package). The model provider is chosen at runtime; Excel is parsed with exceljs (a security choice over SheetJS); the whole thing runs via tsx and, for the cloud, is esbuild-bundled into an arm64 image on AWS AgentCore with scale-to-zero billing.
flowchart LR
F["Raw CSV / Excel"] --> PR["profile"]
PR --> IN["inspect distinct values"]
IN --> M{{"Model reasons a per-column plan"}}
M --> AP["apply plan (fix / flag / leave)"]
AP --> AN["characterize: stats, distributions, ranked correlations"]
AN --> OUT["Ready dataset + cleaning report + pre-analysis report + summaries"]
M -. risky change .-> GATE["Approval gate (held w/ reason)"]
Built With
- amazon-bedrock
- amazon-bedrock-agentcore
- cloudfront
- express.js
- react
- strands-agent
- vite
Log in or sign up for Devpost to join the conversation.