Benchmarks, measurements and explainers on open models, local inference and AI agents, written by Vikas Goenka. Every number comes from a run described in the article. New articles are published here; earlier ones first appeared on ringarc.ai. RSS.
What an AI agent harness is, one real run step by step, the evidence that the harness changes results with the model fixed, and who sells or publishes one in October 2026.
A map of always-on personal AI agents: OpenAI dots, Grok Bot, Meta Muse, Gemini Spark, Microsoft Autopilot, Anthropic's products, OpenClaw and Hermes. Hosted versus self-hosted, who each is for, two worked examples, and how thin the evidence is.
claude-reader is a small TUI that tails the JSONL transcript Claude Code already writes, keeps only the prose and your prompts, and renders them as markdown in a second terminal pane. No hooks, no server, works over ssh. Why the mouse is off, why there is no clock, and what an outside review found before release.
I read the dependency files of 703 starred GitHub repos tagged as AI agents. Raw SDKs beat all frameworks combined, more apps use LangChain's parts than LangChain, and production apps keep frameworks slightly more often than demos do.
Models with million-token context windows still lose facts buried in the middle of a prompt. The research behind context rot: attention sinks, positional decay, distractors, and why the needle-in-a-haystack benchmark on every model card is the easiest possible test.
Qwen3.6-35B-A3B, running 4-bit on a single 16GB RTX 5070 Ti, lands at +0.005 [−0.01, +0.02] on my private 163-task benchmark: statistically at the frontier, tied for the best point estimate in the series, perfect scores on four dimensions, 66 tokens per second, $1.24 for the whole run. Plus why the MoE fits where the smaller dense model doesn't, and two harness fixes worth knowing.
Opus 5 used 1.81x more output tokens for statistically identical quality, showed no measurable capability regression, and refused enough harmless long-context tasks to make that dimension unscoreable. Why a stochastic refusal can quietly manufacture a capability regression on any leaderboard.
There is no RAG debate, there are five: grep vs vectors, long context, GraphRAG, CAG, and memory. Pulling them apart shows why the boring baseline keeps winning, what chunking and indexing actually need, and when the advanced tier earns its complexity.
Three engines, three model sizes, one 16GB GPU, benchmarked head to head: why vLLM collapsed to 33 tok/s then won the whole batch round, how llama.cpp runs an 18GB model at 71 tok/s, and the one env var every Ollama user should set.
Release-day run of Qwen3.8-Max on my private 163-task benchmark: first positive point estimate against frozen frontier anchors, the head-to-head split with Kimi K3, two verified failure stories, a 3x cost gap, and three serving traps to know before you integrate it.
How and why I built project-brain: an open-source, plain-Markdown, cross-project memory for Claude Code. The pain, the research into what exists, the layered architecture, and the convergence with Karpathy's LLM-wiki pattern and Google's OKF.