Inspiration

In modern financial markets, algorithmic trading is heavily bottlenecked by general-purpose software. Even fast code runs through an operating system, language runtimes, and extra software layers between market price updates and order decisions. At the same time, traders and financial professionals who understand strategy logic rarely have the technical background to write low-level code or program FPGAs using languages like Verilog.

We built M4ntis to solve this gap. Instead of forcing users to choose between complex code or software delays, M4ntis lets users build trading strategies through a visual drag and drop interface and run them on a custom CPU designed directly for an FPGA. Rather than writing basic software for a general purpose chip, we designed the chip hardware around the trading problem itself.

What It Does

M4ntis is a platform that converts visual block workflows into real-time hardware execution:

  • Visual Builder and AI Assistant: Users build trading logic using a node interface (similar to Scratch) and can use an integrated LLM to help explain and refine their strategies.
  • Custom 27-Opcode CPU: Strategies compile down to a 27-opcode instruction set running on a custom FPGA CPU designed from scratch, not an off the shelf processor.
  • Native Trading Primitives: The hardware includes single-cycle support for key trading needs like reading price history from circular buffers, calculating moving averages, comparing thresholds, branching, and updating account balances.
  • Fast Deployment: Strategies load over UART in milliseconds without needing hardware resynthesis, letting users iterate quickly on real silicon.
  • Live Dashboard: The FPGA receives live price ticks and streams execution decisions (BUY/SELL) and running balances back to our web interface.

How We Built It

We built M4ntis as a joint software and hardware system:

  • Web Frontend: Built a visual canvas using Next.js and ReactFlow, connected to a Node.js streaming backend and an LLM API to break down strategy structures.
  • Custom ISA Design: Designed an instruction set architecture with dedicated registers, a balance tracking register, five 30-entry circular stock price buffers, and bit-exact instruction encoding.
  • FPGA Hardware: Implemented the CPU in Verilog on a Real Digital Urbana board (Xilinx Spartan-7). We built it in stages: register files and ALU first, then control flow, circular buffers, a multi-cycle divider, balance tracking, and the UART link.
  • Verification: Verified each stage against self-checking testbenches, followed by full Vivado synthesis and place-and-route to verify hardware timing closure and resource usage.

Challenges We Ran Into

  • Timing Closure vs Simulation: Our design passed logical simulation easily, but failed real hardware timing at the board's default 100MHz clock. Simulation only checks logical correctness, so we only caught this by running Vivado synthesis checkpoints. We fixed this by adding an MMCM clock generation block to run a stable 50MHz clock with safe setup timing margins.
  • Full-Stack Scope: Balancing the entire pipeline from web UI, LLM analysis, data normalization, and serial protocols down to custom hardware required careful planning so every part remained fully functional.
  • Environment and Cable Issues: Debugging hardware setup issues took time, including a charge-only USB cable that carried power but no data, missing Vivado cable drivers, and Windows path issues with spaces in folder names.

Accomplishments That We're Proud Of

  • Custom Hardware Execution: Successfully designing, verifying, and closing hardware timing for a custom trading CPU on an FPGA.
  • Automated Test Coverage: Built self-checking testbenches that caught subtle logic bugs early, including a bug where buy and sell balance updates were inverted.
  • Seamless Abstraction: Connecting high-level visual nodes in a web browser straight down to low-level silicon logic without requiring users to write Verilog.

What We Learned

  • Simulation is Not Hardware: Logical simulation passes do not guarantee hardware success. Real timing closure, RAM block inference, and physical I/O behave differently than a logic simulator suggests.
  • Spec Precision Matters: Small questions like whether "balance" means cash on hand or net spend, or how circular buffer boundaries handle overflow, required explicit decisions before writing hardware logic.
  • Hardware/Software Co-Design: Building the web interface and CPU architecture at the same time forced us to carefully plan how compiler logic and hardware execution interact.

What's Next for M4ntis

  • Complete Hardware Bring-Up: Finalizing live UART communication on physical hardware using real tick feeds.
  • Full Strategy Benchmarks: Running full multi-indicator strategies (like a moving average crossover) end-to-end on live price data.
  • Concurrency and Performance: Expanding the CPU architecture to handle multiple stock symbols concurrently and testing maximum throughput now that timing closure is stable.

Built With

Share this project:

Updates

Submission history