Retrieval-Augmented Generation (RAG)
August 26, 2026
Corpus2Skill: Turning Enterprise Knowledge into Navigable Agent Skills
Yiqun Sun, Pengfei Wei, Lawrence B. Hsieh
arXiv

Retrieval-Augmented Generation (RAG) gives language models access to external evidence. A query is embedded, similar passages are retrieved, and the model receives those passages as context for its answer. This simple pattern has become a foundation for enterprise question answering, customer support, and knowledge-grounded assistants.

But conventional RAG treats the model as a passive consumer of search results. The model sees the passages returned by the retriever, but not how the knowledge base is organized, which topics have not yet been explored, or whether a different branch might contain better evidence. Even agentic RAG often amounts to repeated search: the agent can issue more queries, but it still has to guess what to search for without a map of the corpus.

Our new paper introduces Corpus2Skill, a system-level retrieval architecture that changes the interface between an LLM agent and a bounded knowledge base. An offline compiler distills the corpus into a hierarchical directory of informational skills. At serve time, the agent navigates that directory—starting with a bird's-eye view, drilling into progressively finer summaries, retrieving the underlying documents, and backtracking or crossing branches when its first route is unproductive.

The central idea is straightforward: for a well-structured corpus, make knowledge navigable as well as searchable.

‍

From retrieval to navigation

Traditional RAG gives the model a fixed set of top-ranked passages. If the right evidence is missing, the model may not know what else exists or where to look next.

Corpus2Skill instead exposes the organization of the corpus itself. Each top-level skill represents a broad topical partition. Inside it, INDEX.md files describe progressively narrower groups until the agent reaches rows representing individual documents. Lightweight descriptions are visible first; fuller summaries and raw documents are loaded only when the agent chooses to inspect them.

This progressive-disclosure design gives the agent enough information to plan without placing the full knowledge base into its context window. It can keep multiple candidate branches open, compare their coverage, abandon a weak route, follow entity-based cross-links, and combine evidence from different parts of the hierarchy.

The result is not simply a new retriever. It is a compile-then-navigate architecture in which the hierarchy, its cross-links, and the agent's browsing behavior work together as one system.

‍

The offline compiler: distilling a corpus into skills

Corpus2Skill begins with a one-time compilation phase.

Each source document is converted into a compact summary card containing a title, a one-line description, and distinctive phrases. The card is combined with the source text and embedded, giving the clustering process both semantic and surface-level signals.

The compiler then repeatedly clusters related items and asks an LLM to summarize each group. This bottom-up process creates a multi-level topical hierarchy. Once the hierarchy is stable, Corpus2Skill materializes it as a forest of filesystem-based skill directories:

  • each top-level topic becomes a SKILL.md file;
  • intermediate and leaf groups become INDEX.md files;
  • full source documents remain in a separate document store; and
  • an entity index and related-skill links connect information that spans branches.

Three enrichments make the hierarchy more useful to an agent. Soft assignment cross-references documents when they plausibly belong to more than one cluster. Exemplars show representative documents for each branch. The entity index lets the agent jump directly to skills associated with a named product, organization, or feature.

The navigation files remain compact—typically under 2 KB—so the agent can inspect summaries cheaply before loading full documents.

‍

The serve phase: follow the map, then retrieve evidence

At query time, the agent begins with the names and short descriptions of the available skills. This is its bird's-eye view of the corpus.

A typical query takes two or three navigation steps:

  1. The agent selects one or more promising top-level skills.
  2. It reads the relevant SKILL.md and descends into an INDEX.md subgroup.
  3. It retrieves one or more full documents and synthesizes a grounded answer.

If the first branch is unproductive, the agent can backtrack. If a query spans topics, it can follow related-skill links or consult the entity index and combine evidence across branches. The system requires final claims to trace to retrieved source documents rather than to the routing summaries themselves.

This separation is important: summaries help the agent decide where to look, while source documents determine what it can claim.

Corpus2Skill uses ordinary file-navigation and document-lookup tools. It does not require a vector database at serve time, and the compiled directories can be integrated into existing agent environments that support filesystem-based skills.

‍

Evaluation on enterprise question answering

We first evaluate Corpus2Skill on WixQA, an enterprise customer-support benchmark containing 6,221 support articles and 200 expert-written test queries. The compiler produces a three-level hierarchy with six top-level skills and 665 navigation files.

We compare against five baselines spanning different retrieval paradigms:

  • BM25 keyword retrieval;
  • dense vector retrieval;
  • hybrid sparse–dense retrieval;
  • RAPTOR hierarchical retrieval; and
  • an agent with iterative access to sparse, dense, and hybrid search tools.

All systems use the same answer model and evaluation protocol. We measure lexical and semantic answer quality, factuality, context recall and precision, grounded faithfulness, hallucination rate, turns, and per-query cost.

‍

Results: better answers and broader evidence coverage

On WixQA, Corpus2Skill records the highest score on every answer-quality metric and both retrieval-coverage metrics.

Token F1 reaches 0.456, a 21% relative improvement over the agentic retrieval baseline at 0.378 and a 25% improvement over dense retrieval at 0.364. The lead over RAPTOR is statistically significant, and three independent serving runs show only ±0.002 variation.

Factuality reaches 0.767, about six percentage points above the next-best methods. Context recall and precision rise to 0.708 and 0.829, compared with 0.618 and 0.659 for RAPTOR. These gains indicate that the agent is not merely producing more fluent answers—it is locating a more complete and relevant set of supporting documents.

Grounding also improves sharply over unconstrained agentic search. The agentic baseline hallucinates on 50% of WixQA queries under the paper's strict threshold, while Corpus2Skill reduces that rate to 4.5%. Single-shot retrieval remains slightly stronger on some grounding measures because it exposes the model to only a small fixed passage, but it trails substantially in answer quality and evidence coverage.

The trade-off is cost. With prompt caching, Corpus2Skill costs $0.153 per query, roughly 1.9 times the agentic baseline and 13–22 times the single-shot retrievers. A smaller serving model lowers the measured cost to $0.093 while preserving most of the F1 gain, though with weaker grounding. The method is therefore most attractive for high-value questions where completeness and factual support justify additional navigation.

‍

When navigation helps—and when it does not

Corpus navigation is not a universal replacement for retrieval. We evaluate the same system across ten RAGBench subsets, producing an eleven-dataset study together with WixQA. Under paired significance tests, Corpus2Skill wins on five datasets, ties on three, and loses on three.

The pattern follows the shape of the underlying corpus:

  • ‍Navigation works best for bounded, single-domain collections with a recoverable topical taxonomy and documents that each cover a recognizable feature or topic.‍
  • Methods tend to tie on heterogeneous or numerical-table collections where no single organization offers a consistent advantage.‍
  • Flat retrieval remains preferable for open-domain factoid pools, homogeneous collections of near-identical tables, and long documents split across many clause types.

In the failure cases, top-level summaries become generic or nearly indistinguishable, so the hierarchy cannot provide a useful routing signal. This scope condition is a practical design guideline: a corpus should be compiled into skills only when its structure is meaningful enough for an agent to navigate.

‍

What drives the gains

The ablations point to navigability itself rather than model scale as the main source of improvement.

Removing the entity index reduces WixQA F1 from 0.456 to 0.312. Removing soft assignment lowers it to 0.333, and removing branch exemplars lowers it to 0.401. By contrast, replacing the larger compilation encoder with a model 13 times smaller has almost no effect.

The hierarchy also scales predictably. Expanding a corpus from 5,000 to 50,000 documents adds only one navigation level. In the 50,000-document experiment, all queries traverse the larger tree successfully and navigation degrades less than sparse, dense, or hybrid retrieval.

These results suggest that well-designed routing structure can matter more than using the most powerful encoder or allowing the agent many additional search rounds.

‍

Practical considerations

Corpus2Skill is a good fit for mostly stable enterprise collections such as product documentation, policy libraries, technical support centers, and curated internal knowledge bases. Compilation takes approximately 20 minutes for the 6,221-document WixQA corpus, after which the resulting skills can serve many queries.

Frequently changing corpora are more challenging because the current system requires recompilation rather than incremental hierarchy repair. Top-level routing errors also remain the largest source of failure. Future systems could combine navigation and retrieval dynamically—using the hierarchy when the corpus has a strong taxonomy and falling back to flat retrieval for queries or collections that do not.

‍

Looking ahead

Corpus2Skill reframes enterprise RAG as an information-navigation problem. Instead of repeatedly asking a retriever for isolated passages, the agent receives a compact map, decides which branches merit attention, and knows what remains unexplored.

The broader lesson is that an LLM's interface to knowledge matters. For bounded corpora with recoverable structure, exposing that structure can improve answer quality, evidence coverage, and grounding—even without a vector database at serve time.

Code: github.com/dukesun99/Corpus2Skill

‍

‍

‍