Inspiration
Most "multi-agent" systems are sequential pipelines where one AI simply hands work to the next. While effective for simple workflows, they rarely challenge each other, making it easy for mistakes, insecure code, or poor architectural decisions to propagate.
I wanted to build an actual AI society where specialized agents collaborate, critique one another, and improve the final result through structured discussions rather than simple handoffs.
What It Does
APEX Society transforms a single natural language prompt into a production-ready software project using 9 specialized Qwen-powered AI agents.
The workflow consists of:
- Planner
- Architect
- Coder
- Reviewer
- SelfHealer
- Debugger
- Executor
- DocWriter
- Memory
Instead of accepting the first generated solution, the Reviewer evaluates every project using a structured scoring system.
If the reviewer score falls below 80/100, or a critical security issue is detected, the generated code is rejected. The Reviewer sends detailed feedback identifying the affected files and issues. The Coder regenerates only those files before submitting another revision.
This conflict-resolution loop can execute up to three review rounds before continuing through SelfHealer, Debugger, Executor, and documentation generation.
A live dashboard visualizes every interaction between agents in real time, allowing users to observe planning, coding, reviews, debugging, and security improvements as they happen.
During one benchmark:
• Planner and Architect executed in parallel
• Memory injected relevant context from previous projects
• Coder generated 10 production-ready files
• Reviewer approved the submission with a reviewer score of 100/100
• SelfHealer patched all generated files
• Debugger resolved a runtime issue
• Executor validated every generated source file
Compared to a single-agent workflow, APEX Society generated a more complete project while improving review quality and security.
How I Built It
The backend is built using FastAPI and deployed on Alibaba Cloud ECS. All AI inference is performed through Alibaba Cloud DashScope using Qwen 3.7 Max as the primary model.
Each agent has a clearly defined responsibility:
• Planner — project planning and file structure
• Architect — software architecture and security decisions
• Coder — production-ready code generation
• Reviewer — structured code quality and security evaluation
• SelfHealer — automatic vulnerability remediation
• Debugger — runtime error detection and repair
• Executor — syntax validation of generated files
• DocWriter — professional documentation generation
• Memory — retrieves relevant knowledge from previous projects
Planner and Architect execute simultaneously using parallel threads, reducing orchestration latency compared with sequential execution.
The orchestration engine coordinates every interaction, manages Reviewer↔Coder negotiation cycles, and broadcasts all agent events over WebSockets.
The frontend provides:
- Live agent conversations
- Animated SVG agent network
- Quality score updates
- Agent status visualization
- Conflict resolution timeline
Every decision made by an agent is visible in real time.
Challenges
The most difficult engineering challenge involved parsing generated code.
Different Qwen responses produced filenames using different Markdown conventions. Earlier versions of the parser stopped after detecting the first matching format, causing valid generated files to be ignored whenever multiple formatting styles appeared within the same response.
The parser was redesigned to process every code block independently, allowing mixed formatting styles to coexist without losing files.
Another challenge involved synchronizing dashboard animations. Reviewer rejection events completed too quickly to observe visually, so state-locking logic was introduced to preserve rejection animations before transitioning back to normal states.
Accomplishments
- Built a collaborative 9-agent AI software engineering system
- Implemented Reviewer ↔ Coder conflict resolution
- Added parallel Planner and Architect execution
- Integrated cross-session memory retrieval
- Automatic security review and vulnerability remediation
- Live WebSocket dashboard visualizing agent collaboration
- MCP server exposing agent capabilities
- Successfully deployed on Alibaba Cloud ECS
- 35 automated tests passing
What I Learned
Prompt engineering benefits greatly from concrete examples instead of abstract instructions.
Showing examples of secure and insecure implementations consistently produced better code than general rules such as "avoid hardcoded secrets."
Separating responsibilities across specialized agents also produced more reliable software than relying on a single general-purpose model.
What's Next
Future work includes:
- Multimodal project generation directly from wireframes using Qwen 3.7 Plus
- Expanded long-term memory across repositories and projects
- Additional parallel execution of independent agents
- GitHub repository creation and automatic commits
- CI/CD deployment directly to Alibaba Cloud
Built With
- alibabacloudecs
- api
- dashscope
- fastapi
- python
- qwencloud
- svganimatemotion
- websockets
Log in or sign up for Devpost to join the conversation.