Weft 1.0.0
Weft is public. One repo, one CLI, one shared brain: github.com/MennoAf/weft.
Install with pipx install git+https://github.com/MennoAf/weft.git, run weft up, register the MCP server with your agent, and you have persistent memory that every agent — on every machine — can read and write.
What shipped
- A validated snapshot. The release ships from the tested state of the repository: the database-backed suite runs in testcontainers isolation, each test against a fresh Postgres + pgvector instance.
- A benchmark harness with a budget ledger. The LongMemEval runner (
benchmarks/longmemeval/,faithful_s36profile) ingests each question’s haystack through Weft’s normal write path — raw memories and per-turn episode records (dual ingest) — answers under a bounded tool-round policy, and prices every provider call against a conservative reservation ledger before dispatch. The run manifest pins dataset checksum, source hashes, and pricing by hash; the runner refuses drift. - Honest numbers, with provenance. We ran the full LongMemEval-S split (500 questions, turn tier, top_k=10, dual ingest) on 2026-09-29, one authorized paid run. Judged accuracy: 335 / 496 = 67.54% (67.0% against the full selected 500); task-averaged 70.06%. Writer:
gpt-6-luna; judge:gpt-4owith LongMemEval’s official answer-check prompt. Actual spend $0.65 against a $4.79 conservative reservation estimate; wall time ~5.3 hours. It’s a single repetition — evidence, not a publication gate.
What didn’t go well — on purpose
Weft saves what you tell it to save. It doesn’t curate your conversations into benchmark-shaped facts, and the weakest question types show it: multi-session 52.6%, knowledge-update 55.1%, single-session-preference 56.7%. The full per-type table, methodology, cost ledger, and reproduction steps live in docs/benchmarks.md — that document is the source of truth; this note only summarizes it.
Roadmap gates
Falsifiable claims, recorded against commits:
- Multi-session and knowledge-update lift — the two weakest types (52.6%, 55.1%) are the active targets, and the first post-launch work. A retrieval change ships only if these move without regressing the single-session types.
- Turn-tier recall@10 — must lift ≥3 points over baseline; oracle recall must lift ≥1.
When Weft gains an autonomous classifier, it will be because it makes the core product better. Today it saves what you tell it to save — and that is the product.