Inspiration

Every developer has lived this: you join a project, open an issue, and have no idea where to start. You grep for something and get 47 results. You ask a colleague why a technology exists and they say "I think someone decided that two years ago." The knowledge exists — it's in merge requests, issues, commit messages, and the import graph. GitLab Orbit indexes all of it into a queryable knowledge graph. GitLab Knowledge Navigator is the translation layer that turns natural developer questions into Orbit graph queries and surfaces actionable answers.

What it does

GitLab Knowledge Navigator gives every developer a Staff Engineer in their terminal — one that knows the entire codebase, remembers every decision, and never goes on vacation.

Before Navigator, a developer spends hours tracing files, merge requests, and ownership information manually. With Navigator, onboarding, impact analysis, architecture understanding, and repository risk discovery become single-command workflows grounded in Orbit's knowledge graph. It also generates a self-contained executive HTML dashboard combining repository scoring, drift detection, alerts, impact analysis, interactive knowledge graphs, and AI-synthesized recommendations — shareable with anyone, no install required.

The tool ships 10 commands:

  • navigator onboard --issue N traces an issue through Orbit's knowledge graph to find relevant files, identify human maintainers, and suggest a concrete first change
  • navigator explain --component X pulls architecture, MR history, dependents, and owners from Orbit and synthesizes a complete component overview
  • navigator why --query "why does httpx exist here?" traces technology adoption through import chains and MR history with a dated chronological timeline — returning answers like "httpx was introduced June 19, 2026 via MR !4 as an async HTTP client for Orbit REST API queries"
  • navigator impact --file X traverses Orbit's dependency graph, categorizes dependents by architectural layer, flags untested files, and classifies risk as Low / Medium / High / Critical
  • navigator drift --project X detects circular dependencies, dependency hotspots, and high-fanout modules — surfaced as a named architectural risk report
  • navigator alerts --project X fires proactive risk alerts when two or more signals intersect: no MR history, no documentation, no test coverage, high dependent count
  • navigator executive --project X synthesizes dependency concentration analysis and score data into leadership-level risk summaries in plain business language
  • navigator compare --query httpx runs the same question with and without Orbit data side by side — making the value of graph-grounded AI immediately visible
  • navigator dashboard --output report.html generates the unified self-contained HTML report with all of the above
  • navigator demo --issue 1 runs the complete new-contributor journey in sequence

The tool works on any GitLab repository in your namespace. It was validated on both the Knowledge Navigator project itself and a fork of gitlab-org/gitlab-runner — a 1,366-file Go codebase with 9,314 definitions and 58,340 indexed relationships — demonstrating multi-repository intelligence at real production scale.

How we built it

Knowledge Navigator uses a hybrid Local + Remote Orbit architecture — the key engineering insight of the project.

Orbit Remote (/api/v4/orbit/query) provides SDLC context: WorkItems, MergeRequests, Project members, and SDLC relationships. We use traversal, search, and neighbors query types across 5 entity types.

Orbit Local (glab orbit local CLI + DuckDB) provides source-code structure: Files, Definitions, ImportedSymbols, and dependency edges via SQL. This enables multi-hop graph walks like "find every file that imports a class that depends on OrbitClient" — queries the REST API cannot replicate.

The hybrid design was built directly from hitting Orbit's remote indexing latency during development — rather than silently returning empty results, the tool detects incomplete remote indexing and automatically falls back to the local graph with a single informational banner. Multi-repo support uses automatic clone + local index: resolve_local_path() clones any GitLab project to ~/.navigator/repos/, indexes it with glab orbit local, and scopes all SQL queries to that project by ID to prevent cross-project data contamination.

AI synthesis is handled by Google Gemini 2.5 Flash via the google-genai SDK. Every agent passes real Orbit graph data to Gemini as structured context — the model synthesizes from facts, not from hallucination.

The project includes 22 passing pytest tests across 5 test files covering score formula math, alert trigger logic, risk classification, client initialization, and SQL query construction. The CI pipeline runs the full test suite on every push.

Challenges we ran into

Orbit indexing latency was the biggest engineering challenge. The remote code indexer for our test project returned 0 File/Definition counts for over 14 hours despite healthy infrastructure (32/32 replicas ready). This forced the hybrid architecture decision — not as a fallback plan, but as a core design requirement. The local DuckDB graph became the primary source for all source-code analysis, with Remote Orbit handling SDLC metadata it indexes reliably and fast.

Cross-project data isolation in the local DuckDB graph required careful SQL scoping by project ID throughout every query. The local Orbit graph stores all indexed repositories in a single DuckDB file. Without explicit project-ID scoping, running analysis on a second repository silently blended data from both projects, producing incorrect dependency counts and false architectural findings. Every SQL query in the tool is scoped to the specific project being analyzed.

The hybrid architecture trade-off meant designing two different query paths — SQL for local graph traversal, REST DSL for remote SDLC data — and ensuring both paths produce consistent, comparable results. The has_test_coverage(), has_docstring(), and find_file_dependents() helpers are shared across the Impact, Executive, and Alerts agents to guarantee cross-feature data consistency.

Accomplishments that we're proud of

Validated on real production-scale data. Running navigator drift on the gitlab-runner fork — a 1,366-file Go codebase with 9,314 definitions and 58,340 indexed relationships — returned common/config.go with 27 dependents classified Critical, with real Go file names and real dependency counts. This proves the tool works at production scale, not just on the small project it lives in.

Cross-feature data consistency. When Impact Analysis reports 7 dependents for navigator/orbit/client.py, the Executive Report also reports 7, and the Alerts agent also reports 7 — because all three call the same find_file_dependents() shared helper. This consistency was caught and fixed mid-build: the Executive agent had its own raw SQL aggregation that included self-loops, producing an off-by-one vs Impact Analysis. The fix unified all three under one function, guaranteeing a single version of the truth.

The "Why Orbit?" contrast. The side-by-side comparison between a generic LLM answer for "why does httpx exist here?" (generic async theory) and the Orbit-grounded answer ("introduced June 19 via MR !4 for async HTTP to the Orbit REST API") is the single clearest demonstration of what graph-grounded AI makes possible. No prompt engineering creates this — it requires the graph data that only Orbit provides.

What we learned

Orbit Local is dramatically underused. The remote API gets all the attention, but the local DuckDB graph enables SQL-level dependency traversal that the REST API cannot replicate. Multi-hop walks like "find all files that transitively depend on client.py" require the local graph's edge table. The hybrid architecture unlocked capabilities impossible with Remote-only integration.

Real data matters more than feature count. The seed MRs created with descriptive titles ("feat: introduce httpx as async HTTP client for Orbit REST API queries") transformed the Decision Timeline feature from returning empty results to returning a compelling dated answer. The feature didn't change — the data did. A tool is only as good as the graph it queries.

Indexing latency is a product constraint, not a bug. Designing around it — transparently, with graceful fallback — produced a more robust architecture than assuming the remote graph would always be current.

What's next for GitLab Knowledge Navigator

Integration tests with recorded HTTP fixtures — the current test suite covers logic with mocks; the natural next step is integration tests against recorded Orbit API responses that can run in CI without live credentials.

Streaming output — long-running commands like navigator dashboard currently batch all output at the end. Streaming each section as it completes would significantly improve the interactive experience on large repositories.

Scheduled intelligence reports — a GitLab CI template that runs navigator score and navigator alerts on a schedule and posts results as pipeline artifacts or MR comments, turning the tool from on-demand to continuously proactive.

Support for GitLab Groups — currently scoped to individual projects; extending to group-level analysis would enable cross-project dependency mapping and organization-wide intelligence scoring.

GitLab Knowledge Navigator demonstrates that GitLab Orbit is more than a search API — it can serve as the intelligence layer for developer-focused agents that understand architecture, history, ownership, and impact across an entire software project.

Built With

  • ci/cd
  • duckdb
  • gitlab
  • gitlab-orbit-(remote-api-+-local-duckdb)
  • glab-orbit-local-cli
  • google-gemini-2.5-flash
  • google-genai-sdk
  • pytest
  • python
  • rich-terminal-ui
  • vanilla-svg-+-javascript
Share this project:

Updates

Submission history