AchiralAchiral
Cognoscenti

Cognoscenti: A Benchmark for Trustworthy AI Memory

Cognoscenti evaluates whether AI memory can be trusted in use: whether the right context activates, stale context stays quiet, and repeated experience becomes easier to recall without becoming indiscriminate storage.

It is designed for cognitive memory systems and Emergent Memory Systems where memory is shaped by salience, recency, reinforcement, decay, and context rather than simple persistence or vector similarity.

Why trustworthy memory needs its own benchmark

AI memory does not fail only by forgetting. It also fails when it recalls obsolete instructions, activates plausible but wrong context, treats every stored fact as equal, or cannot explain why one memory was selected over another. Cognoscenti makes those behaviors measurable.

What Cognoscenti Measures

Activation

Tests whether the current query, project, account, or workflow brings the right memory forward first.

Selective forgetting

Tests whether stale, superseded, low-signal, or misleading memories stay out of the active context.

Interference resistance

Tests whether related but wrong memories can be suppressed when they compete with the correct context.

Consolidation

Tests whether repeated and validated experience becomes easier to retrieve without turning everything into permanent storage.

Adaptation

Tests whether new facts update older assumptions instead of leaving the system anchored to stale context.

Efficiency

Tests whether useful memory can remain selective as the memory set grows and distractors accumulate.

Current Public Run

The first public organizational-memory run, dated 2026-08-02, compares RAG, agent-memory, and organic-memory reference baselines on the same inspectable JSONL workload. Organic memory led the tested architectures on Top-1 accuracy and distractor suppression while matching agent memory on Recall@3.

Top-1 accuracy

91.67%

Organic-memory reference baseline.

Recall@3

100%

Matched by agent memory and organic memory.

Distractor activation

16.67%

Lowest rate among the tested references.

Workload

12

Auditable organizational-memory items.

Reproducible Artifacts

The public run does not claim to benchmark Mem0, Zep, Letta, Glean, LangGraph, or any other vendor product. Vendor results should be published only when provider endpoints, keys, commands, seeds, and raw outputs are recorded.

FAQ

What is Cognoscenti?

Cognoscenti is Achiral's benchmark framework for trustworthy AI memory. It evaluates whether a memory system activates the right context, suppresses stale or misleading context, adapts through use, and remains auditable.

What makes an AI memory benchmark trustworthy?

Trustworthy AI memory requires more than recall. A benchmark should measure activation, forgetting, interference resistance, consolidation, adaptation, and whether raw artifacts are available for inspection.

Is Cognoscenti only for Achiral?

No. Cognoscenti includes local reference baselines and supports external memory-system evaluation through HTTP adapters. Vendor-specific results should be published only when endpoints, keys, seeds, commands, and raw outputs are recorded.

Read the current benchmark report.

See how Cognoscenti evaluates memory behavior under use.

Open benchmark report