CANOPY: Adaptive-Granularity Evidence
Compression for Multimodal RAG

Hyojeong Yun1,*, Jueun Kim2,*, Wook-Shin Han1,†

1 GSAI, POSTECH  ·  2 CSE, POSTECH

* Equal contribution  ·  † Corresponding author

Overview of CANOPY: each retrieved item becomes a hierarchy of original regions, parent-relative refinement selects regions at different depths, and a critic guides additional retrieval

CANOPY projects each retrieved item into a hierarchy of its own regions, keeps a region only where it stands out against its parent, and asks a critic whether what has been kept is enough to answer.

Abstract

Multimodal RAG retrieves text, tables, images, and videos, but choosing a retrieval granularity does not determine how much context to retain within each item. Coarse units include irrelevant content, while uniformly fine selection can remove context needed to interpret the evidence. Existing compressors address this trade-off with modality-specific mechanisms, leaving open a shared procedure for adapting the retained extent region by region across heterogeneous items.

We introduce CANOPY (Canonical Projection over Hierarchy), a framework for adaptive-granularity post-retrieval evidence compression. CANOPY represents retrieved items as hierarchies and uses a node encoder fine-tuned on gold evidence to score regions against the query. Parent-relative refinement compares these scores to select multiple regions at different granularities without LLM calls for node-level pruning. Because compression cannot recover evidence that was never retrieved, a critic requests targeted follow-up retrieval when it judges the accumulated evidence insufficient; newly retrieved items are compressed before being added. Across five QA benchmarks over a 33M-item heterogeneous corpus, CANOPY achieves higher average answer accuracy than the evaluated retrieval baselines. Ablations indicate that additional retrieval drives the main accuracy gains on multi-hop QA. In the unrouted Qwen3-VL-8B-Instruct setting, compression reduces reader-input evidence tokens by 14.2–27.7% relative to the same iterative pipeline without compression, with comparable answer accuracy.

Why within-item granularity

The evidence a question needs is usually a small part of a retrieved item: a sentence in a passage, a row in a table, a segment of a long video. Passing the whole item spends the reader's context on distractors; cutting every item to its finest pieces separates evidence from the context that makes it interpretable. Neither a coarse nor a fine unit is right for every region of every item.

Three retrieved items, a passage, a table with a linked passage, and a video, with the question-relevant evidence highlighted
The evidence is a small part of the item. Blue marks what the question actually needs; CANOPY keeps a sentence (a), a table row and a linked sentence (b), or a video segment (c).

Granularity is decided inside the item

That an item is relevant says nothing about which of its regions the reader needs or how much surrounding context each one requires — and these differ between regions of the same item. So the retained extent has to be decided after retrieval, region by region, not fixed per item by a query-level router.

A comparison, not a threshold or a count

Refinement descends into every child that scores at least as high as its parent and keeps the parent otherwise. The scores come from a node encoder fine-tuned on gold-span preferences. Nothing else is tuned: no absolute similarity threshold, no retained-unit count, and no LLM call for pruning a node.

Compression and retrieval do different jobs

The critic goes and gets what is missing, which is where the accuracy comes from on multi-hop questions. Compression keeps what accumulates small: 14.2–27.7% fewer reader-input tokens than the same iterative pipeline without it, at comparable accuracy.

Method

Hierarchy construction

Each retrieved item becomes a hierarchy of its own regions: the whole item at the root and minimal fragments at the leaves — sentences for text, rows for tables with the column headers kept at every level, 30-second segments for video. Images stay single nodes.

Parent-relative refinement

A frozen query encoder and a node encoder fine-tuned on gold-span preferences score every node against the query. From the root, a node is replaced by every child that scores at least as high as it, and kept whole when no child does. Branches stop independently, so the selected regions form an evidence forest at different depths, carrying original content rather than summaries — and no threshold, count, or LLM call decides a node.

Critic-guided additional retrieval

Compression cannot recover evidence that was never retrieved. After each round a critic checks the accumulated evidence against the original question; if it is insufficient, the critic names the missing fact and issues a follow-up query, whose results are compressed before they are added, for at most three rounds.

Results

Retrieval runs over a unified corpus of about 33M items built from the collections of NQ, HotpotQA, OTT-QA, MMQA, and LVBench: 32.4M passages, 429K tables, 57K images, and 1.5K video segments. The reader is Qwen3-VL-8B-Instruct, k = 10 items per round, at most three rounds. Evidence is retrieved, never given.

NQ HotpotQA OTT-QA MMQA LVBench
Method EMF1EMF1EMF1EMF1Acc.Avg.
No Retrieval17.529.023.331.47.511.923.127.130.420.4
Vanilla RAG single round38.252.638.949.712.917.540.145.934.432.9
UniversalRAG routed35.050.436.546.710.514.436.942.041.832.1
IRCoT iterative35.051.347.461.824.330.838.946.917.832.7
CANOPY39.754.449.662.127.733.445.151.135.039.4
CANOPY routed38.152.747.660.120.124.942.147.443.438.3
Exact match / F1 on the four text-based benchmarks and accuracy on multiple-choice LVBench, with Qwen3-VL-8B-Instruct as the reader. Avg. is the mean of EM (Acc. for LVBench). Routed variants send each query to modality-specific corpora before retrieval, following UniversalRAG. Evidence tokens, latency, LLM calls, and the InternVL3.5-8B reader are in the paper.
39.4 avg. EM / Acc. best of the six methods over the five benchmarks — 6.5 above Vanilla RAG and 6.7 above IRCoT
14.2 – 27.7% fewer tokens reader-input evidence, against the same iterative pipeline without compression, at comparable accuracy
2.8 LLM calls / question against 5.6 for S2G-RAG at comparable F1 — pruning inside an item costs no LLM call
Pooled answer quality versus reader-input evidence tokens for complete retrieval pipelines at k = 5, 10, 20
Quality against evidence tokens, pooled over the five benchmarks. Each curve runs through retrieval sizes k = 5, 10, 20. CANOPY sits above and to the left of the other complete pipelines: more answer quality for fewer tokens passed to the reader.

Retrieval brings the accuracy; compression bounds the tokens

Removing iterative retrieval lowers HotpotQA and OTT-QA exact match by 10.1 and 14.9 points: on multi-hop questions the second piece of evidence becomes identifiable only after the first has been read, and a single round cannot fetch it. Removing compression instead leaves accuracy where it is — every EM change falls inside its 95% confidence interval — but raises reader-input tokens by 14.2–27.7%, because each round then adds whole items rather than the regions that were retained. Two rounds on average cost CANOPY at most 28% more tokens than single-round Vanilla RAG, and 46–88% fewer than UniversalRAG.

Against compression methods built for one modality

Answer quality versus evidence tokens for RECOMP, LongLLMLingua, S2G-RAG, AKS, and CANOPY across NQ, HotpotQA, OTT-QA, and LVBench
Each method at several evidence budgets. Lines connect a method's budget settings in increasing order; numbers give the mean LLM calls per question; hollow markers are LLM-based. Larger budgets do not close the gap for LongLLMLingua or AKS. S2G-RAG reaches comparable F1 with fewer tokens but twice the LLM calls.

What fine-tuning the encoder changes

Measured directly on the parent–child decisions of test gold items, the fine-tuned encoder makes the right call 58.5% of the time against 36.9% for the frozen one. Downstream, the same refinement procedure descends below the root for 85–96% of items instead of 30–40%, passing 6.8–19.1% fewer reader-input tokens while F1 and accuracy stay within their confidence intervals.

Median root-relative node scores at each tree level for gold-overlapping and other nodes, with the frozen and the fine-tuned encoder, on text, tables, and video
Root-relative node scores by tree level. With the frozen encoder, gold-overlapping and other regions are hard to tell apart; after fine-tuning the gold regions rise above their parents and the others fall below, on text, tables, and video alike.

CANOPY in action

Three test questions, one per modality, replayed from the run: a gold item as it was retrieved, and what CANOPY kept of it. Press Replay refinement to walk the hierarchy top-down and watch each parent–child comparison decide what the reader sees.

BibTeX

@misc{yun2026canopy,
  title         = {{CANOPY}: Adaptive-Granularity Evidence Compression
                   for Multimodal {RAG}},
  author        = {Yun, Hyojeong and Kim, Jueun and Han, Wook-Shin},
  year          = {2026},
  eprint        = {2610.00923},
  archivePrefix = {arXiv},
  primaryClass  = {cs.IR}
}