What Inspired Us Automation delivery in enterprises is fragmented - requirements live in one document, technical designs get rewritten from scratch, platform resources are provisioned manually, and verification evidence is scattered across terminals and Slack threads. When scope changes or a handoff happens, the entire cycle repeats and critical context is lost.

We experienced this pain firsthand while building UiPath automations. Every project felt like starting from zero, even when the patterns were similar. The breakthrough came when we realized: what if a specialized AI agent could handle each phase of the delivery lifecycle, supervised by a coordinator that ensures nothing falls through the cracks?

That's when UiPlan was born - an agent-of-agents orchestrator that treats automation delivery as a deterministic pipeline, where each phase has clear inputs, outputs, and specialist ownership.

What We Learned Building UiPlan taught us three critical lessons:

  1. Constraints are the real bottleneck, not code generation

Early iterations focused on generating perfect workflow code. But we quickly learned that the hard part isn't generating code - it's enforcing the constraints that keep automations safe and maintainable. We built a constraint intelligence system that:

Extracts rules from three sources (brief, skills, codebase) Classifies severity using pattern matching: $\text{severity} = f(\text{keywords}) \in {\text{high}, \text{medium}, \text{low}}$ Blocks progression on violations before they become production incidents

  1. Evidence-first design beats "trust me" handoffs

Traditional automation projects end with "It works on my machine." UiPlan flips this - every phase produces machine-readable evidence:

Planning contract: spec.md, plan.md, tasks.md Platform verification: actual CLI output from queue/asset checks Execution logs: timestamped command traces The handoff package is the artifact, not the conversation.

  1. Agent specialization > general-purpose prompting

We started with one "builder agent" that did everything. It was terrible. Breaking the pipeline into 7 specialist agents (intake analyst, solution architect, workflow generator, platform provisioner, run verifier) gave each agent a focused job with clear success criteria. The LangGraph state machine ensures they coordinate without stepping on each other.

How We Built It Architecture:

Business Brief (JSON) ↓ LangGraph Orchestrator (Python) ↓ 7-Phase Pipeline: 1. assign_agents → Role mapping 2. generate_design_docs → UiPlan + PDD/SDD/ADD 3. generate_uipath_artifacts → Flow JSON + scripts 4. provision_resources → Queue/asset verification 5. execute_flow → Run generated workflow 6. emit_ui_events → Stream to viewer 7. summarize_handoff → Package deliverables ↓ Complete Delivery Package Tech Stack:

LangGraph: State machine orchestration with typed edges UiPath Python SDK: Platform integration (queues, assets, jobs) UiPath CLI (uip): Resource provisioning and monitoring CopilotKit: Interactive viewer with 7 monitoring tabs Pydantic: Type-safe state management and validation Key Implementation Details:

State Graph Pattern: Each phase is a pure function: (OrchestratorState) → OrchestratorState. No side effects in phase logic - all state changes are explicit.

Constraint Extraction: Regex-based pattern matching for critical rules:

CONSTRAINT_PATTERN = r"\b(must|never|required|do not|should not|blocker)\b" Classified by severity using keyword analysis.

Evidence Collection: Every CLI invocation is logged with timestamps and exit codes. The execution-evidence.json file is the single source of truth for "did it work?"

Event Streaming: The emit_ui_events phase serializes the entire orchestrator state to JSON for the viewer. The viewer polls this file for updates (future: WebSocket streaming).

Challenges We Faced Challenge 1: LangGraph learning curve

LangGraph's checkpoint-based state management was powerful but unfamiliar. We struggled with understanding when state mutations persisted vs. when they were lost.

Solution: Adopted a strict pattern - every phase returns a new dict with only the keys it modified. The graph merges these updates automatically.

Challenge 2: CLI output parsing is brittle

The uip CLI outputs vary by version and command. JSON mode helps, but some commands don't support it.

Solution: Built robust error handling - if JSON parsing fails, fall back to regex extraction. If that fails, log the raw output and escalate to human review. Better to admit uncertainty than fail silently.

Challenge 3: Constraint dependency cycles

Some constraints reference each other: "Never deploy to Production" depends on "Must verify folder name". Naive topological sort fails on cycles.

Solution: Instead of enforcing a strict ordering, we classify constraints by severity and validate all high-severity constraints before each phase. Cycles are allowed - we just ensure critical rules are checked repeatedly.

Challenge 4: Real-time viewer updates

Early versions used WebSockets for live updates. Deploying this to static GitHub Pages broke everything.

Solution: Simplified to file polling. The orchestrator writes run-events.json, the viewer polls it every 2 seconds. Works everywhere, zero infrastructure.

Challenge 5: Making generated docs actually useful

First iterations generated generic templates. No one read them.

Solution: Two changes -

Templates now have {{PLACEHOLDER}} markers that must be filled with actual data The viewer shows a checklist: "✓ Constraints propagated to SDD", "✓ Success criteria linked to test plan" This forced us to make generation deterministic and reviewable.

What's Next If we continue this project:

WebSocket streaming for real-time viewer updates Multi-run comparison - diff two runs side-by-side in the viewer Approval gates - human-in-the-loop checkpoints at phase boundaries Constraint learning - extract new constraints from failed runs automatically

Built With

  • apis
  • cli
  • copilotkit
  • npmn
  • python
  • uipath
Share this project:

Updates