<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>ProbeArc tech blog</title>
  <link>https://probearc.ai/blog</link>
  <atom:link href="https://probearc.ai/blog/rss.xml" rel="self" type="application/rss+xml"/>
  <description>Benchmarks, measurements and explainers on open models, local inference and AI agents.</description>
  <language>en</language>
  <lastBuildDate>Thu, 08 Oct 2026 13:30:21 +0000</lastBuildDate>
  <item>
    <title>AI agent harnesses demystified: what they are, how one works, and who offers them</title>
    <link>https://probearc.ai/blog/agent-harnesses-explained</link>
    <guid isPermaLink="true">https://probearc.ai/blog/agent-harnesses-explained</guid>
    <pubDate>Thu, 08 Oct 2026 09:00:00 +0000</pubDate>
    <description>What an AI agent harness is, one real run step by step, the evidence that the harness changes results with the model fixed, and who sells or publishes one in October 2026.</description>
  </item>
  <item>
    <title>Personal AI agents in October 2026: hosting choices and what we know</title>
    <link>https://probearc.ai/blog/personal-agents-compared</link>
    <guid isPermaLink="true">https://probearc.ai/blog/personal-agents-compared</guid>
    <pubDate>Mon, 05 Oct 2026 09:00:00 +0000</pubDate>
    <description>A map of always-on personal AI agents: OpenAI dots, Grok Bot, Meta Muse, Gemini Spark, Microsoft Autopilot, Anthropic&#x27;s products, OpenClaw and Hermes. Hosted versus self-hosted, who each is for, two worked examples, and how thin the evidence is.</description>
  </item>
  <item>
    <title>Tail the log: a terminal reading pane for Claude Code sessions</title>
    <link>https://probearc.ai/blog/claude-reader</link>
    <guid isPermaLink="true">https://probearc.ai/blog/claude-reader</guid>
    <pubDate>Mon, 17 Aug 2026 09:00:00 +0000</pubDate>
    <description>claude-reader is a small TUI that tails the JSONL transcript Claude Code already writes, keeps only the prose and your prompts, and renders them as markdown in a second terminal pane. No hooks, no server, works over ssh. Why the mouse is off, why there is no clock, and what an outside review found before release.</description>
  </item>
  <item>
    <title>I checked what popular agent software actually uses</title>
    <link>https://probearc.ai/blog/what-agent-apps-actually-use</link>
    <guid isPermaLink="true">https://probearc.ai/blog/what-agent-apps-actually-use</guid>
    <pubDate>Sat, 15 Aug 2026 09:00:00 +0000</pubDate>
    <description>I read the dependency files of 703 starred GitHub repos tagged as AI agents. Raw SDKs beat all frameworks combined, more apps use LangChain&#x27;s parts than LangChain, and production apps keep frameworks slightly more often than demos do.</description>
  </item>
  <item>
    <title>Context Rot: Why AI Models Lose Track of Long Prompts</title>
    <link>https://probearc.ai/blog/context-rot-lost-middle</link>
    <guid isPermaLink="true">https://probearc.ai/blog/context-rot-lost-middle</guid>
    <pubDate>Fri, 14 Aug 2026 09:00:00 +0000</pubDate>
    <description>Models with million-token context windows still lose facts buried in the middle of a prompt. The research behind context rot: attention sinks, positional decay, distractors, and why the needle-in-a-haystack benchmark on every model card is the easiest possible test.</description>
  </item>
  <item>
    <title>Frontier quality now runs on a 16GB gaming GPU</title>
    <link>https://probearc.ai/blog/qwen36-35b-local-frontier</link>
    <guid isPermaLink="true">https://probearc.ai/blog/qwen36-35b-local-frontier</guid>
    <pubDate>Thu, 13 Aug 2026 09:00:00 +0000</pubDate>
    <description>Qwen3.6-35B-A3B, running 4-bit on a single 16GB RTX 5070 Ti, lands at +0.005 [−0.01, +0.02] on my private 163-task benchmark: statistically at the frontier, tied for the best point estimate in the series, perfect scores on four dimensions, 66 tokens per second, $1.24 for the whole run. Plus why the MoE fits where the smaller dense model doesn&#x27;t, and two harness fixes worth knowing.</description>
  </item>
  <item>
    <title>I Tested the Complaints About Opus 5. One Was True. One Wasn&#x27;t. One I Couldn&#x27;t Test.</title>
    <link>https://probearc.ai/blog/opus5-complaints-measured</link>
    <guid isPermaLink="true">https://probearc.ai/blog/opus5-complaints-measured</guid>
    <pubDate>Wed, 12 Aug 2026 09:00:00 +0000</pubDate>
    <description>Opus 5 used 1.81x more output tokens for statistically identical quality, showed no measurable capability regression, and refused enough harmless long-context tasks to make that dimension unscoreable. Why a stochastic refusal can quietly manufacture a capability regression on any leaderboard.</description>
  </item>
  <item>
    <title>Everyone Tells You Basic RAG Is Dumb. It Is Not!</title>
    <link>https://probearc.ai/blog/rag-five-debates</link>
    <guid isPermaLink="true">https://probearc.ai/blog/rag-five-debates</guid>
    <pubDate>Mon, 10 Aug 2026 09:00:00 +0000</pubDate>
    <description>There is no RAG debate, there are five: grep vs vectors, long context, GraphRAG, CAG, and memory. Pulling them apart shows why the boring baseline keeps winning, what chunking and indexing actually need, and when the advanced tier earns its complexity.</description>
  </item>
  <item>
    <title>Ollama vs llama.cpp vs vLLM on one 16GB desktop card</title>
    <link>https://probearc.ai/blog/local-inference-comparison</link>
    <guid isPermaLink="true">https://probearc.ai/blog/local-inference-comparison</guid>
    <pubDate>Fri, 07 Aug 2026 09:00:00 +0000</pubDate>
    <description>Three engines, three model sizes, one 16GB GPU, benchmarked head to head: why vLLM collapsed to 33 tok/s then won the whole batch round, how llama.cpp runs an 18GB model at 71 tok/s, and the one env var every Ollama user should set.</description>
  </item>
  <item>
    <title>Same tier, different personalities: Qwen3.8-Max vs Kimi K3 on my private benchmark</title>
    <link>https://probearc.ai/blog/qwen38-max-vs-kimi-k3</link>
    <guid isPermaLink="true">https://probearc.ai/blog/qwen38-max-vs-kimi-k3</guid>
    <pubDate>Fri, 07 Aug 2026 09:00:00 +0000</pubDate>
    <description>Release-day run of Qwen3.8-Max on my private 163-task benchmark: first positive point estimate against frozen frontier anchors, the head-to-head split with Kimi K3, two verified failure stories, a 3x cost gap, and three serving traps to know before you integrate it.</description>
  </item>
  <item>
    <title>Everyone&#x27;s giving Claude a brain. I just wanted to stop repeating myself.</title>
    <link>https://probearc.ai/blog/project-brain</link>
    <guid isPermaLink="true">https://probearc.ai/blog/project-brain</guid>
    <pubDate>Fri, 07 Aug 2026 09:00:00 +0000</pubDate>
    <description>How and why I built project-brain: an open-source, plain-Markdown, cross-project memory for Claude Code. The pain, the research into what exists, the layered architecture, and the convergence with Karpathy&#x27;s LLM-wiki pattern and Google&#x27;s OKF.</description>
  </item>
</channel>
</rss>
