<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Ruben Marcus — Blog</title><description>Technical articles on AI agents, benchmarks, WebGL, and the stack behind rubenmarcus.dev.</description><link>https://rubenmarcus.dev/</link><language>en-us</language><item><title>From prompt to product: five ways to build with AI</title><link>https://rubenmarcus.dev/blog/from-prompt-to-product-five-ways-to-build-with-ai/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/from-prompt-to-product-five-ways-to-build-with-ai/</guid><description>Prompt engineering, vibe coding, agentic engineering, product engineering, and research engineering sound like names for the same thing. They are not. The difference is what you control, what feedback you trust, and who decides the work is done.</description><pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate><category>ai</category><category>agents</category><category>product-engineering</category><category>research</category></item><item><title>My AI harness for frontend: from prompt to pull request</title><link>https://rubenmarcus.dev/blog/frontend-ai-harness-prompt-to-pull-request/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/frontend-ai-harness-prompt-to-pull-request/</guid><description>The system I use to turn an idea into verifiable frontend: specs, skills, model selection, Ralph Starter, worktrees, browsers, screenshots, tests, GitHub, and telemetry. The model writes code. The harness decides what to trust.</description><pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate><category>ai</category><category>agents</category><category>frontend</category><category>harness</category><category>ralph-starter</category></item><item><title>Inside the Gauntlet loop</title><link>https://rubenmarcus.dev/blog/inside-the-gauntlet-loop/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/inside-the-gauntlet-loop/</guid><description>The techniques inside the adversarial agent loop that builds CS Brasil: a 25-criterion visual rubric written so a language model can grade a PNG, critic prompts that ban vague answers by name, a generated symbol-level conflict table for a 6,543-line file, and a regression hunter with permission to find nothing.</description><pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate><category>ai</category><category>agents</category><category>gamedev</category><category>orchestration</category></item><item><title>This portfolio is agents-welcome. Probably the first.</title><link>https://rubenmarcus.dev/blog/agents-welcome-portfolio/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/agents-welcome-portfolio/</guid><description>My site has an AGENTS.md, an MCP server, a hiring API, and a terminal resume. Your agent can read my CV, check my availability, and book an intro. Here is how it works.</description><pubDate>Thu, 06 Aug 2026 00:00:00 GMT</pubDate><category>ai</category><category>agents</category><category>mcp</category><category>portfolio</category></item><item><title>The AI harness behind the CS Brasil game</title><link>https://rubenmarcus.dev/blog/cs-brasil-ai-harness/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/cs-brasil-ai-harness/</guid><description>The machine that builds CS Brasil is no longer a folder of markdown. It is a measurement engine: the real game booted in pure Node, 61 invariants with mutation tests, generated docs that fail CI on drift, and four laws learned the expensive way.</description><pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate><category>ai</category><category>agents</category><category>gamedev</category><category>harness</category></item><item><title>A command center for agent swarms, in markdown</title><link>https://rubenmarcus.dev/blog/agent-command-center/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/agent-command-center/</guid><description>I run a swarm of coding agents — Codex, Claude, Amp, Kimi — plus humans, coordinated entirely through markdown files. No database, no dashboard-as-source-of-truth, chat history explicitly banned. Here is the shared brain: control room, task queue, candidate cemetery, and fail-closed automation.</description><pubDate>Sun, 02 Aug 2026 00:00:00 GMT</pubDate><category>ai</category><category>agents</category><category>orchestration</category><category>tooling</category></item><item><title>I built my portfolio with a fleet of AI agents</title><link>https://rubenmarcus.dev/blog/i-built-my-portfolio-with-a-fleet-of-ai-agents/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/i-built-my-portfolio-with-a-fleet-of-ai-agents/</guid><description>The making-of of this site: AI-generated 3D models that kept coming out as busts, a shader pipeline that morphs me into wireframe and ASCII, a visual gauntlet that grades every pixel, and what I learned about art-directing machines.</description><pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate><category>ai</category><category>agents</category><category>threejs</category><category>webgl</category><category>making-of</category></item><item><title>How AEO can help your business grow</title><link>https://rubenmarcus.dev/blog/aeo-what-it-moves/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/aeo-what-it-moves/</guid><description>AEO is the plumbing that lets ChatGPT, Perplexity and friends read and cite your business. I built a free scanner that has run 4,569 scans across 2,259 unique sites; the same 3 failures show up everywhere.</description><pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate><category>aeo</category><category>seo</category><category>llm</category><category>web</category></item><item><title>Building a browser FPS with AI agents</title><link>https://rubenmarcus.dev/blog/shipping-a-browser-fps/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/shipping-a-browser-fps/</guid><description>CS Brasil is a satirical browser FPS — 5 factions, 44 characters, 26 weapons, 5 maps, zero-build Three.js — built almost entirely by AI agents running a named adversarial loop: the Gauntlet. Critics grade, builders edit the same 6,543-line file in parallel, and the loop never ends itself.</description><pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate><category>gamedev</category><category>ai-agents</category><category>webgl</category><category>javascript</category><category>threejs</category></item><item><title>Streaming 2.85M messages: the plumbing of a production agent chat</title><link>https://rubenmarcus.dev/blog/vercel-ai-sdk-streaming/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/vercel-ai-sdk-streaming/</guid><description>I was the second most active committer on the open-source chat package of a production agent platform, the streaming front of an AI runtime that turned OpenAPI specs into tools and drove on-chain agents across NEAR and EVM. Real code from BitteProtocol/chat: the useChat loop, tool-call states, the OpenAPI operationId contract, and how history gets rebuilt.</description><pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate><category>ai</category><category>vercel</category><category>streaming</category><category>nextjs</category></item><item><title>I rebuilt my agent loop in Mastra. Here&apos;s what my runtime gets right.</title><link>https://rubenmarcus.dev/blog/mastra-field-notes/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/mastra-field-notes/</guid><description>I ran hand-rolled agent loops in production on a web3 agent platform (2.85M+ messages) and in ralph-starter. I spent a week rebuilding the same loop in Mastra to see what a framework buys. Head-to-head notes: what Mastra makes easy, what my runtime does that it can&apos;t, and who should pick which.</description><pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate><category>ai</category><category>agents</category><category>mastra</category><category>typescript</category></item><item><title>Routing 9 agent roles across 7 providers: the ECDSA.fail harness</title><link>https://rubenmarcus.dev/blog/openrouter-routing/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/openrouter-routing/</guid><description>I led AI engineering on the multi-agent research harness that took #1 on ECDSA.fail. The routing layer is the part worth stealing: role to model tables, fail-closed adapters, and a spend gate that aborts the run before the bill does.</description><pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate><category>ai</category><category>llm</category><category>openrouter</category><category>tooling</category></item><item><title>Cross-pollinating LLMs: peer review for machines</title><link>https://rubenmarcus.dev/blog/llm-cross-pollination/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/llm-cross-pollination/</guid><description>I built a research pipeline where frontier models from different providers generate, adversarially review, and merge each other&apos;s work. The only thing shared between them is the previous phase&apos;s text. It turns out peer review compiles into prompt choreography surprisingly well.</description><pubDate>Sat, 11 Jul 2026 00:00:00 GMT</pubDate><category>ai</category><category>llm</category><category>agents</category><category>research</category><category>openrouter</category></item><item><title>Keeping an autonomous research agent honest</title><link>https://rubenmarcus.dev/blog/autoresearcher-pareto-frontier/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/autoresearcher-pareto-frontier/</guid><description>Autoresearcher runs benchmark-driven research loops: one agent step, one benchmark command, one number, keep or reject. The part that makes it work isn&apos;t the agent. It&apos;s that the loop is deliberately simple, fail-closed, and leaves a full audit trail in git and JSONL.</description><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate><category>ai</category><category>agents</category><category>evaluation</category><category>benchmarks</category><category>open-source</category></item><item><title>Evals are the product</title><link>https://rubenmarcus.dev/blog/evals-are-the-product/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/evals-are-the-product/</guid><description>What an eval actually is, learned the expensive way: a browser FPS gated by 61 invariants that must each ship with a failing mutation, and a quantum decoder benchmark won with one scalar and a five-step loop. The model changes every few months. The ruler survives.</description><pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate><category>ai</category><category>evals</category><category>agents</category><category>testing</category><category>llm</category></item><item><title>Git worktrees are my agent orchestrator</title><link>https://rubenmarcus.dev/blog/dag-agent-orchestration/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/dag-agent-orchestration/</guid><description>ralph-starter runs coding agents in Ralph Wiggum loops, and swarm mode coordinates them with git worktrees, Promise.allSettled, and a strategy switch. Three strategies with real tradeoffs, and why a DAG executor would reject this design at startup.</description><pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate><category>ai</category><category>agents</category><category>orchestration</category><category>typescript</category></item><item><title>The swarm that took #1 on ECDSA.fail</title><link>https://rubenmarcus.dev/blog/the-agent-swarm-that-took-1-on-ecdsa-fail/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/the-agent-swarm-that-took-1-on-ecdsa-fail/</guid><description>How I built an autonomous multi-agent research harness — 9 specialist LLM roles across 7+ providers, fail-closed adapters, lane-isolated dispatch, and a falsifier queue — that produced the top-ranked quantum circuit on ECDSA.fail.</description><pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate><category>ai</category><category>agents</category><category>orchestration</category><category>benchmarks</category><category>quantum-computing</category></item><item><title>Context engineering inside a runtime with 344K chats</title><link>https://rubenmarcus.dev/blog/context-engineering/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/context-engineering/</guid><description>A production AI-agent runtime I helped build served 2.85M messages across 344K chats with streaming agent loops and OpenAPI to tool conversion. The context was the spec: system prompts lived in x-mb, tools were operationIds, and wallet state rode every request. Real code from BitteProtocol/chat, make-agent, and agent-sdk.</description><pubDate>Fri, 05 Jun 2026 00:00:00 GMT</pubDate><category>ai</category><category>agents</category><category>context</category><category>llm</category></item><item><title>How RAG works inside Mirofi.sh</title><link>https://rubenmarcus.dev/blog/rag-in-production/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/rag-in-production/</guid><description>Mirofi.sh is a hosted multi-agent social-simulation platform built on an open-source engine, GraphRAG, and Zep. This is the retrieval story: what agent memory stores, how it feeds each turn, and where the graph earned its complexity.</description><pubDate>Thu, 28 May 2026 00:00:00 GMT</pubDate><category>ai</category><category>rag</category><category>llm</category><category>production</category></item><item><title>The Mini Shai-Hulud Case and the Real Risk of Dependencies</title><link>https://rubenmarcus.dev/blog/mini-shai-hulud-dependency-risk/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/mini-shai-hulud-dependency-risk/</guid><description>The May 2026 Mini Shai-Hulud wave didn&apos;t just steal an npm token — it abused the build and publishing pipeline itself: pull_request_target, cache poisoning, OIDC trusted publishing, and install scripts. What actually happened, and what I&apos;d audit first.</description><pubDate>Tue, 12 May 2026 00:00:00 GMT</pubDate><category>security</category><category>supply-chain</category><category>npm</category><category>ci-cd</category><category>dependencies</category></item><item><title>How I Hit #1 on a Quantum Error Correction Challenge using AI</title><link>https://rubenmarcus.dev/blog/how-i-hit-1-qec-using-ai/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/how-i-hit-1-qec-using-ai/</guid><description>I topped the QEC Decoder Optimization Arena leaderboard at 2,642 errors per million — not with physics credentials, but with a tight experimental loop: held-out validation, statistical gating, and a swarm of agents in isolated git worktrees.</description><pubDate>Thu, 16 Apr 2026 00:00:00 GMT</pubDate><category>quantum-computing</category><category>ai</category><category>agents</category><category>benchmarks</category><category>optimization</category></item><item><title>Automating entire workflows with ralph-starter</title><link>https://rubenmarcus.dev/blog/automating-entire-workflows-with-ralph-starter/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/automating-entire-workflows-with-ralph-starter/</guid><description>ralph-starter runs Ralph Wiggum loops — fetch a spec, run the AI agent, check tests/lint/build, feed errors back, repeat. Here&apos;s how it works and why I built it.</description><pubDate>Thu, 19 Feb 2026 00:00:00 GMT</pubDate><category>ai</category><category>automation</category><category>ralph-wiggum</category><category>open-source</category></item><item><title>Getting started with Next.js + Strapi: Security first</title><link>https://rubenmarcus.dev/blog/getting-started-with-next-js-strapi-security-first/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/getting-started-with-next-js-strapi-security-first/</guid><description>Before touching Content-Types or routes, talk about security. This is the security playbook for a Strapi + Next.js stack — XSS, CSRF, clickjacking, JWT, SSL, monitoring.</description><pubDate>Sun, 16 May 2021 00:00:00 GMT</pubDate><category>nextjs</category><category>frontend</category><category>strapi</category><category>security</category></item><item><title>Why use Next.js + Strapi?</title><link>https://rubenmarcus.dev/blog/why-use-next-js-strapi/</link><guid isPermaLink="true">https://rubenmarcus.dev/blog/why-use-next-js-strapi/</guid><description>The case for pairing Next.js with the Strapi headless CMS for a modern decoupled frontend / CMS stack. Headless concepts, alternatives, boilerplates.</description><pubDate>Fri, 07 May 2021 00:00:00 GMT</pubDate><category>nextjs</category><category>react</category><category>strapi</category><category>headless</category></item></channel></rss>