Release notes

What changed, when, and why — newest first. Numbers and claims link back to the repo, which stays the source of truth. Follow along with RSS.

Weft 1.0.0

launchbenchmarksroadmap

Weft is public. One repo, one CLI, one shared brain: github.com/MennoAf/weft.

Install with pipx install git+https://github.com/MennoAf/weft.git, run weft up, register the MCP server with your agent, and you have persistent memory that every agent — on every machine — can read and write.

What shipped

  • A validated snapshot. The release ships from the tested state of the repository: the database-backed suite runs in testcontainers isolation, each test against a fresh Postgres + pgvector instance.
  • A benchmark harness with a budget ledger. The LongMemEval runner (benchmarks/longmemeval/, faithful_s36 profile) ingests each question’s haystack through Weft’s normal write path — raw memories and per-turn episode records (dual ingest) — answers under a bounded tool-round policy, and prices every provider call against a conservative reservation ledger before dispatch. The run manifest pins dataset checksum, source hashes, and pricing by hash; the runner refuses drift.
  • Honest numbers, with provenance. We ran the full LongMemEval-S split (500 questions, turn tier, top_k=10, dual ingest) on 2026-09-29, one authorized paid run. Judged accuracy: 335 / 496 = 67.54% (67.0% against the full selected 500); task-averaged 70.06%. Writer: gpt-6-luna; judge: gpt-4o with LongMemEval’s official answer-check prompt. Actual spend $0.65 against a $4.79 conservative reservation estimate; wall time ~5.3 hours. It’s a single repetition — evidence, not a publication gate.

What didn’t go well — on purpose

Weft saves what you tell it to save. It doesn’t curate your conversations into benchmark-shaped facts, and the weakest question types show it: multi-session 52.6%, knowledge-update 55.1%, single-session-preference 56.7%. The full per-type table, methodology, cost ledger, and reproduction steps live in docs/benchmarks.md — that document is the source of truth; this note only summarizes it.

Roadmap gates

Falsifiable claims, recorded against commits:

  • Multi-session and knowledge-update lift — the two weakest types (52.6%, 55.1%) are the active targets, and the first post-launch work. A retrieval change ships only if these move without regressing the single-session types.
  • Turn-tier recall@10 — must lift ≥3 points over baseline; oracle recall must lift ≥1.

When Weft gains an autonomous classifier, it will be because it makes the core product better. Today it saves what you tell it to save — and that is the product.