← Index

MetaNaviT: Retrieval over Research Files

Find the current configuration, trace the supporting files, and review proposed changes before applying them.

FIG. 05 — Illustrated retrieval workflow on demo files. This is not a large-corpus or live-model benchmark.

Problem

Research folders contain active configs, archived drafts, logs, and checkpoints. Finding a plausible answer is easier than identifying the current file and returning traceable evidence.

Approach

Combined BM25 and dense retrieval with reciprocal rank fusion, query routing, staleness filtering, and typed graph traversal. Exposed search and proposed filesystem changes through an MCP interface.

What I built

In a six-person team, I worked on retrieval, the data layer, and APIs. The repository combines hybrid retrieval, evaluation fixtures, and approval-gated file operations. The lightweight benchmark uses explicit hash embeddings and overlap reranking rather than downloaded embedding models.

Result

The committed 136-query fixture benchmark reports Recall@50 0.938 and nDCG@10 0.493 using hash embeddings and overlap reranking. On eight exact-path queries, adding routing after reranking increased nDCG@10 from 0.875 to 0.938.

Follow the answer back to a file

The walkthrough starts with conflicting copies of a configuration. Retrieval surfaces the current source and its location; the user can inspect that evidence instead of relying on an unsupported answer.

FIG. 05A — Illustrated ranking pipeline: lexical and dense candidates merge through RRF.

Make proposed changes reviewable

Mutation tools return a plan before application. The MCP walkthrough shows a blocked change and an approved change. The next benchmark step is a real corpus with actual embedding and cross-encoder backends, evaluated separately from these fixtures.

FIG. 05B — Demo workflow: propose, review, approve, apply.

What the benchmark supports

The frozen CPU fixture contains 136 queries over 61 files. BM25 scores 0.505 nDCG@10 versus 0.493 for the hybrid configuration; the benchmark does not establish an overall ranking improvement or production-scale retrieval quality.

Stack

Links