Case
studies
Two problems I went deep on at Axy — from the initial problem to the technical approach and what we shipped.
Safe, reviewable edits to a live knowledge graph
A scientific knowledge graph is only useful if researchers can trust the data in it. We needed a system that let researchers propose graph changes — add nodes, update edges, merge entities — without those changes going live until a reviewer approved them. And every change needed a complete audit trail.
Researchers needed to draft graph changes collaboratively, get them reviewed, and merge them into the main graph without corrupting live data. Naive approaches — writing directly to the graph and rolling back on rejection — were too risky. We also needed conflict detection: what happens when two researchers edit the same node?
I designed a delta-based submission model. Rather than writing directly to the graph, each submission stores a structured diff — the intended changes — alongside metadata: author, timestamp, affected entities, and a conflict fingerprint. During review, a conflict detection pass compares the submission's target state against the current graph. If another submission has already changed the same entities, the conflict surfaces before merge, not after.
Merging a submission is handled as an async job: changes are applied atomically, the mutation history is appended, and downstream services (search indexes, vector store, document store) are notified via SQS. Dead-letter queues and reconciliation checks ensure consistency even when a downstream service is temporarily unavailable.
Researchers can draft and submit graph changes through a structured UI. Reviewers see a clear diff, any detected conflicts, and the mutation history for the affected entities. Approved submissions merge cleanly without touching the live graph during review. The delta model also made rollback straightforward: reverting a merge is just applying the inverse diff.
From baseline retrieval to research-grade search
Scientific search is harder than general-purpose search. Queries use domain-specific terminology, abbreviations, and entity names that sparse keyword search handles well but dense semantic search misses — and vice versa. We needed both working together.
Our initial retrieval setup was evaluated on scientific BEIR benchmarks (Recall@K, NDCG@10, MRR). It was an initial baseline — usable, but not good enough for a research tool where missing a relevant entity could meaningfully affect a scientist's work. The challenge was building a retrieval pipeline that could handle both exact-term queries and conceptual, semantic queries.
I built a multi-stage retrieval pipeline in Qdrant. The first stage runs dense vector search (embedding-based) and sparse keyword search simultaneously. Both result sets are then fused using Reciprocal Rank Fusion (RRF) — a rank-based fusion method that doesn't require calibrating score scales across different retrieval methods.
The fused candidates go through a reranking step that scores candidate-query relevance more precisely. I also added query rewriting to expand ambiguous or abbreviated scientific queries before retrieval, which helped on queries where the user's phrasing didn't match the document's phrasing. Evaluation was run on scientific BEIR subsets to track progress across iterations.
Retrieval quality improved significantly across Recall@K, NDCG@10, and MRR on scientific BEIR evaluations. The pipeline also feeds the RAG system — better retrieval meant better context for the AI pipeline, which improved answer quality downstream. The evaluation framework I set up made it possible to measure the effect of each change, so future improvements could be benchmarked rather than guessed at.
More Case Studies Coming Soon
Currently documenting upcoming breakdowns on distributed event queues with AWS SQS/SNS, multi-database synchronization patterns, and internal tooling architecture.