North Star
Weekly Active Namespaces
142
namespaces with 1+ query / 7 days
+38 vs prior 30 days
Our single most important number. A namespace is "active" if at least one person or agent ran a GKG query in the past 7 days. Example: if JPMC's engineers ran 50 queries this week, that's 1 active namespace.
Invocation Method
Where queries originate
gkg_query_executed - source_type (30d)
Duo Agent Platform (DAP)
61%
External Agent via MCP
24%
REST API (direct)
10%
CLI (glab)
5%
DAP (Duo Agent Platform)
Zero-rated. Included in Duo Developer. Orbit agent within Duo responds to natural language questions using GKG as its knowledge layer. Example: "What pipelines are failing in my org?" answered by Duo, powered by GKG.
External Agent via MCP (Claude Code, Codex)
Billed via GitLab Credits. External agents connect using
query_graph and get_graph_schema MCP tools. This is our monetization signal. Example: Claude Code asking "what service owns this file?" during a code review.Usage
Total Queries (30d)
48,291
+22% vs prior period
Daily Active Users
317
+14% vs prior period
Namespaces Indexed
289
of 312 enabled
23 pending or failed
Avg Queries / Session
6.4
DAP sessions only
up from 3.1 last month
Namespaces by Query Volume
gkg_query_executed - namespace_id grouped by query count bucket (30d)
Power users (1k+ queries)
4
namespaces
Light users (1-10 queries)
62
namespaces — at risk of churning
How many namespaces fall into each usage tier. Most are still light users (1-10 queries/month), which is expected for an early beta. The goal over time: move namespaces up and to the right. A namespace running 200+ queries/month is deeply integrated. Source: COUNT(*) of gkg_query_executed grouped by namespace_id, then bucketed.
Query Volume by Type
gkg_query_executed - query_type
The four ways you can ask GKG a question. find_nodes = look up specific things ("show me all services"). traverse = follow relationships ("what calls this function?"). explore = browse a subgraph. aggregate = count or summarize ("how many pipelines failed this week?"). High traverse volume signals agents are doing real dependency analysis.
Top SDLC Entities Queried
gkg_query_executed - entity_types_queried
What data people are actually asking GKG about. High pipeline volume = customers using GKG for CI/CD health and inheritance mapping. High definitions volume = code graph use cases like blast radius analysis. This tells us which use cases are winning in the wild.
Query Performance
Query Result Status
gkg_query_executed - result_status
⚠
18% empty result rate. Queries returning 0 rows are a silent failure. Top pattern: traversal queries on un-indexed TypeScript repos.
Success = query ran and returned results. Empty = query ran but found nothing (the dangerous one: the agent gets no data and has no idea why). Error = something broke. Empty results are worse than errors because they're silent: the agent gets no context, leading to worse AI responses with no indication something went wrong.
Query Latency Breakdown by Type
gkg_query_executed - compile / authorization / execute duration (p50, ms)
Authorization (Rails)
Compile (SQL)
Execute (ClickHouse)
Every query has three steps: check permissions (Authorization), convert the request to SQL (Compile), then run it against the database (Execute). The purple bar dominates because GKG has to verify what namespaces you can access on every single query. If that bar shrinks, it means Rails auth caching is working. traverse queries take longest because they follow chains of relationships across the graph.
ClickHouse Data Read per Query Type
gkg_query_executed - ch_read_bytes (avg MB)
How much raw data ClickHouse has to scan to answer each query type. traverse reads 18MB on average because following relationships requires chaining multiple table joins. This is our primary cost driver: high MB = high ClickHouse compute cost. Target is to bring traverse below 10MB via materialized views or pre-computed paths.
Indexing Health
Indexing Success Rate by Language
gkg_indexing_completed - status + languages_indexed
When a namespace enables GKG, we parse and index their code. This shows how reliably that works per language. Partial = some files were skipped (usually due to stack overflows on large nested expressions). TypeScript at 63% is our biggest risk: many enterprise customers use it heavily.
| Language | Success Rate | Avg Files Indexed | Status | |
|---|---|---|---|---|
| Ruby | 98% | 12,400 | Stable | |
| Java | 94% | 8,100 | Stable | |
| Kotlin | 91% | 3,200 | Stable | |
| Python | 87% | 5,800 | Watch | |
| TypeScript | 63% | 4,100 | At Risk | |
| JavaScript | 71% | 2,900 | Watch |
Graph Coverage per Namespace
gkg_indexing_completed - avg node counts
Avg Definitions / NS
43,200
Avg Edges / NS
218,000
The "size" of the knowledge graph we built for a typical namespace. Definitions = functions, classes, and methods we indexed (43K avg means a medium-large codebase). Edges = the relationships between them: calls, imports, inherits. More edges = richer graph = better answers. A namespace with 218K edges can answer "what calls this?" across the whole org.
Avg Index Duration
8.4s
initial full index
Avg Watermark Lag
2.1m
SDLC data freshness
within SLO
Partial Index Rate
14%
of all indexing runs
mostly TypeScript / JS
Activation Funnel
Namespace Activation
gkg_namespace_enabled + gkg_indexing_completed + gkg_query_executed
The journey from "admin turned on GKG" to "team uses it every week." Each step is a potential drop-off point we can improve. The biggest gap today is between indexed and first query: 64 namespaces got a full graph built but nobody has run a single query yet. These are activation targets.
GKG Enabled
100%
Indexed Successfully
93%
First Query Run
72%
Active (7d WAN)
46%
Recurring (3+ weeks)
29%
💡
Biggest drop: indexed to first query (93% to 72%). 64 namespaces indexed but never queried. Target for activation nudge or admin UX improvement.
DAP Integration
GKG Calls per DAP Session
gkg_dap_session_summary - total_gkg_calls distribution
How many times an agent called GKG within a single Duo session. 1-2 calls = the agent barely used GKG (treated it as optional). 7+ calls = the agent is deeply relying on the graph to answer the user's question. We want this histogram to shift right over time. The 16+ bucket is agents doing multi-hop analysis: "find service A, traverse to its dependencies, check their pipeline health."
DAP Session Outcomes
gkg_dap_session_summary
Avg GKG Calls / Session
6.4
up from 3.1
Fallback to REST
11%
of sessions
target: below 5%
Avg session tokens: GKG-heavy (8+ calls) vs GKG-light (1-2 calls)
Sessions that use GKG heavily actually consume fewer tokens because the graph returns structured, precise data instead of the agent having to read raw files or call multiple REST endpoints. Lower token usage = faster responses + lower cost per session. This is the core efficiency argument for GKG.