probearc

Tech blog

Benchmarks, measurements and explainers on open models, local inference and AI agents, written by Vikas Goenka. Every number comes from a run described in the article. New articles are published here; earlier ones first appeared on ringarc.ai. RSS.

8 October 2026 · Essay

AI agent harnesses demystified: what they are, how one works, and who offers them

What an AI agent harness is, one real run step by step, the evidence that the harness changes results with the model fixed, and who sells or publishes one in October 2026.

5 October 2026 · Essay

Personal AI agents in October 2026: hosting choices and what we know

A map of always-on personal AI agents: OpenAI dots, Grok Bot, Meta Muse, Gemini Spark, Microsoft Autopilot, Anthropic's products, OpenClaw and Hermes. Hosted versus self-hosted, who each is for, two worked examples, and how thin the evidence is.

17 August 2026 · Tool

Tail the log: a terminal reading pane for Claude Code sessions

claude-reader is a small TUI that tails the JSONL transcript Claude Code already writes, keeps only the prose and your prompts, and renders them as markdown in a second terminal pane. No hooks, no server, works over ssh. Why the mouse is off, why there is no clock, and what an outside review found before release.

15 August 2026 · Study

I checked what popular agent software actually uses

I read the dependency files of 703 starred GitHub repos tagged as AI agents. Raw SDKs beat all frameworks combined, more apps use LangChain's parts than LangChain, and production apps keep frameworks slightly more often than demos do.

14 August 2026 · Explainer

Context Rot: Why AI Models Lose Track of Long Prompts

Models with million-token context windows still lose facts buried in the middle of a prompt. The research behind context rot: attention sinks, positional decay, distractors, and why the needle-in-a-haystack benchmark on every model card is the easiest possible test.

13 August 2026 · Benchmark

Frontier quality now runs on a 16GB gaming GPU

Qwen3.6-35B-A3B, running 4-bit on a single 16GB RTX 5070 Ti, lands at +0.005 [−0.01, +0.02] on my private 163-task benchmark: statistically at the frontier, tied for the best point estimate in the series, perfect scores on four dimensions, 66 tokens per second, $1.24 for the whole run. Plus why the MoE fits where the smaller dense model doesn't, and two harness fixes worth knowing.

12 August 2026 · Benchmark

I Tested the Complaints About Opus 5. One Was True. One Wasn't. One I Couldn't Test.

Opus 5 used 1.81x more output tokens for statistically identical quality, showed no measurable capability regression, and refused enough harmless long-context tasks to make that dimension unscoreable. Why a stochastic refusal can quietly manufacture a capability regression on any leaderboard.

10 August 2026 · RAG series, part 1

Everyone Tells You Basic RAG Is Dumb. It Is Not!

There is no RAG debate, there are five: grep vs vectors, long context, GraphRAG, CAG, and memory. Pulling them apart shows why the boring baseline keeps winning, what chunking and indexing actually need, and when the advanced tier earns its complexity.

7 August 2026 · Benchmark

Ollama vs llama.cpp vs vLLM on one 16GB desktop card

Three engines, three model sizes, one 16GB GPU, benchmarked head to head: why vLLM collapsed to 33 tok/s then won the whole batch round, how llama.cpp runs an 18GB model at 71 tok/s, and the one env var every Ollama user should set.

7 August 2026 · Benchmark

Same tier, different personalities: Qwen3.8-Max vs Kimi K3 on my private benchmark

Release-day run of Qwen3.8-Max on my private 163-task benchmark: first positive point estimate against frozen frontier anchors, the head-to-head split with Kimi K3, two verified failure stories, a 3x cost gap, and three serving traps to know before you integrate it.

7 August 2026 · Open source

Everyone's giving Claude a brain. I just wanted to stop repeating myself.

How and why I built project-brain: an open-source, plain-Markdown, cross-project memory for Claude Code. The pain, the research into what exists, the layered architecture, and the convergence with Karpathy's LLM-wiki pattern and Google's OKF.