Paper results · 125 tasks

InMind
Leaderboard.

A single view of every backbone and query-time memory configuration reported in the paper's main results table. Rankings use end-to-end Application accuracy.

Reference control

Backbone control (in-context)

The target memory is already visible; no retriever or embedding is used.

Target recall100.0%
Application84.0%

Query-time configurations

Leaderboard

Results are ordered by end-to-end Application. Higher is better for every metric, tied scores share a rank, and evidence badges link to the public paper and code where available.

RankMemory system (embedding)Direct recallNaive queryTarget recallIndirect contextApplicationEnd-to-end
01Naive RAG (text-embedding-3-large)Retrieval control97.6%6.4%16.0%
02MemoryOS (text-embedding-3-large)Memory system96.8%7.2%14.4%
03Naive RAG (MiniLM)Retrieval control92.0%0.8%9.6%
03Naive RAG (BM25)Retrieval control53.6%3.2%9.6%
03A-Mem (text-embedding-3-large)Memory system100.0%12.0%9.6%
06HippoRAG 2 (MiniLM)Memory system93.6%0.8%8.8%
07MemoryOS (MiniLM)Memory system87.2%2.4%8.0%
08A-RAG (text-embedding-3-large)Memory system
Paper Code Venue unconfirmed Team verified
93.6%11.2%7.2%
09xMemory (text-embedding-3-large)Memory system
Paper Code Venue unconfirmed Team verified
91.2%6.4%6.4%
09Mem0 (text-embedding-3-large)Memory system76.8%6.4%6.4%
09A-Mem (MiniLM)Memory system99.2%2.4%6.4%
12HippoRAG 2 (text-embedding-3-large)Memory system96.0%2.4%5.6%
13A-RAG (MiniLM)Memory system
Paper Code Venue unconfirmed Team verified
97.6%5.6%4.8%
13xMemory (MiniLM)Memory system
Paper Code Venue unconfirmed Team verified
84.8%3.2%4.8%
15Mem0 (MiniLM)Memory system76.0%2.4%3.2%
01

Direct recall

Can the system retrieve the stored personal fact when the query explicitly asks for it?

02

Target recall

Does the decisive fact reach GPT-5-mini's context for the indirect query?

03

Application

Does the final answer correctly apply the personal fact through the required world-knowledge bridge?

MiniLM denotes all-MiniLM-L6-v2 (384 dimensions); text-embedding-3-large uses 3,072 dimensions. BM25 is lexical retrieval. Scores are percentages on the same 125-task evaluation set.

Verification policy

Unverified scores are reported, not guaranteed.

For entries without an InMind team verification badge, we cannot independently confirm the authenticity of the results. Those scores come from the submitting author team's report and are provided for reference only.

Submit a result

Need a spot on the leaderboard?

Contact the InMind team with your system, configuration, and reproducible result package.

imlrz@mail.ustc.edu.cn