Activation
Tests whether the current query, project, account, or workflow brings the right memory forward first.
Cognoscenti evaluates whether AI memory can be trusted in use: whether the right context activates, stale context stays quiet, and repeated experience becomes easier to recall without becoming indiscriminate storage.
It is designed for cognitive memory systems and Emergent Memory Systems where memory is shaped by salience, recency, reinforcement, decay, and context rather than simple persistence or vector similarity.
AI memory does not fail only by forgetting. It also fails when it recalls obsolete instructions, activates plausible but wrong context, treats every stored fact as equal, or cannot explain why one memory was selected over another. Cognoscenti makes those behaviors measurable.
Tests whether the current query, project, account, or workflow brings the right memory forward first.
Tests whether stale, superseded, low-signal, or misleading memories stay out of the active context.
Tests whether related but wrong memories can be suppressed when they compete with the correct context.
Tests whether repeated and validated experience becomes easier to retrieve without turning everything into permanent storage.
Tests whether new facts update older assumptions instead of leaving the system anchored to stale context.
Tests whether useful memory can remain selective as the memory set grows and distractors accumulate.
The first public organizational-memory run, dated 2026-08-02, compares RAG, agent-memory, and organic-memory reference baselines on the same inspectable JSONL workload. Organic memory led the tested architectures on Top-1 accuracy and distractor suppression while matching agent memory on Recall@3.
Top-1 accuracy
91.67%
Organic-memory reference baseline.
Recall@3
100%
Matched by agent memory and organic memory.
Distractor activation
16.67%
Lowest rate among the tested references.
Workload
12
Auditable organizational-memory items.
Benchmark report
The current public Cognoscenti run comparing RAG, agent memory, and organic memory reference baselines.
Methodology
The scoring rules, workload structure, caveats, and interpretation notes for the public run.
JSONL workload
The inspectable workload with queries, gold memories, distractors, context, and rationale.
Machine-readable results
Raw item-level and aggregate outputs for reproducible review.
The public run does not claim to benchmark Mem0, Zep, Letta, Glean, LangGraph, or any other vendor product. Vendor results should be published only when provider endpoints, keys, commands, seeds, and raw outputs are recorded.
Cognoscenti is Achiral's benchmark framework for trustworthy AI memory. It evaluates whether a memory system activates the right context, suppresses stale or misleading context, adapts through use, and remains auditable.
Trustworthy AI memory requires more than recall. A benchmark should measure activation, forgetting, interference resistance, consolidation, adaptation, and whether raw artifacts are available for inspection.
No. Cognoscenti includes local reference baselines and supports external memory-system evaluation through HTTP adapters. Vendor-specific results should be published only when endpoints, keys, seeds, commands, and raw outputs are recorded.
See how Cognoscenti evaluates memory behavior under use.