The four benchmarks with enough published figures to be worth drawing. One dot per system, every dot a figure as published by the source it cites. Nothing here was measured by this harness.
A solid dot is a project reporting its own score; a faded dot is a score someone else ran. Where a project publishes several figures, the highest is shown. The shaded band marks where the leading group sits, so you can see how tightly the field is packed.
It has no published answer-quality figure on LongMemEval, LoCoMo or any BEAM split. Its only measured results are retrieval recall, which is a different metric family and cannot be put on these axes. Until an answer-quality run exists, it does not appear below.