AVAILABLE FOR FULL-TIME ROLES OPEN TO FREELANCE CONTRACTS AI FULLSTACK ENGINEER · AGENTIC AI AGENT-READY · HUMANS WELCOME AI AGENTS · LLM TOOLING · EVALS NEXT.JS · TYPESCRIPT · REACT BASED IN LISBON · REMOTE WORLDWIDE
← All writing

Tupi: rebuilding 1554 in the browser with 32 AI agents

A 1554 Tupinambá canoe raid rebuilt in real-time 3D in the browser with Claude Opus 5.5, three.js and Blender: the pipeline, what broke, what it cost.

Cover image for Tupi: rebuilding 1554 in the browser with 32 AI agents

Tupi is a three-and-a-half-minute film that runs live in your browser. A Tupinambá war fleet of bark canoes paddles up the Bertioga channel at dawn, around 1554, towards a Tupiniquim village. It loops: near the end it starts to drizzle, thunder rolls over the Serra do Mar, and the next dawn clears it again.

I didn’t write it. I directed it, in one long conversation with Claude Opus 5.5 in Claude Code, which wrote the research dossier, the three.js code, the Blender scripts and the review tooling, and ran 31 other agents to do the parallel work. This post is about how that went: the pipeline, the parts that fought back, and the bill.

A woodcut from Hans Staden's 1557 book next to a frame of the film

The prompt

It started with one message, translated here from Portuguese:

Clone dgreenheck/tidewater. He made a really good visual experiment with Opus 5.5 and I want something at the same level, but in pre-colonial Brazil: a Tupinambá going up a river into enemy territory, set in Pindorama. The idea of this folder is to make several demos portraying historical moments of humanity with three.js and 3D. Do you think we’re capable not only of producing this level of 3D visuals in the browser, but also of doing the correct historical research?

Everything after that was direction in rounds, the way you’d review a film cut:

  • “the trees on the right bank are over-lit”;
  • “the mist looks like smoke and kills the forest’s colour”;
  • “everything with rocks is very low poly”;
  • “were the Serra do Mar mountains really that pointy?”;
  • “add crabs and mangrove animals”.

The full sequence of requests is on the demos page: the ✦ prompt button on the Tupi card.

Research before geometry

The first thing I asked was whether we could hit the visual bar of dgreenheck’s tidewater and Meng To’s Sakura River Valley and get the history right. So the first deliverable wasn’t a scene, it was a cited dossier: Hans Staden (1557), Jean de Léry (1578), Gabriel Soares de Sousa (1587), Eduardo Navarro’s Old Tupi dictionary, Florestan Fernandes, Carneiro da Cunha and Viveiros de Castro.

Every visual choice traces back to it:

  • Canoes. The igaras are bark canoes, not dugouts. Staden describes the bark taken off a tree in one piece, heated over a fire, bent up at both ends and held open by two crossbars, with room for thirty men.
  • Body paint. Genipap (black) and urucum (red).
  • Hair. A tonsure: the front of the crown shaved, the hair left long at the back.
  • Ornaments. The enduape, a rosette of rhea feathers tied at the small of the back, and the tembetá lip plug.
  • Wildlife. Scarlet ibises crossing the channel and mullet jumping, because Staden says the raids came in August, “at the time they go after a kind of fish”.

It also corrected my brief twice. I’d said “Pindorama”, which turns out to be a 19th-century coinage. And the Serra do Mar in the first renders was a row of alpine peaks. When I asked “were the mountains really that pointy?”, the answer was no: it’s an old, eroded escarpment, rounded and covered in rainforest. Here is the round-4 serra next to the final one:

Round 4 with pointy peaks and an amber glare artefact on the water, next to the final rounded Serra do Mar

The whole dossier is public, in Portuguese, with a confidence level on every claim: research.md. If you want to read the sources yourself:

The stack

  • Engine. three.js r186 on WebGPU with TSL node materials, with a WebGL2 fallback.
  • Modules. 14 of them: sky, water, terrain, mangrove, forest, fleet, birds, fauna, smoke, audio, camera, ambience, overlay and post. Each has one owner and the same contract: create(ctx) returns { update(t, dt) }. A module that throws is logged and skipped, so one broken part never blanks the film.
  • Post. TRAA, custom god rays, bloom, auto exposure, a grade and film grain.
  • Camera. It orbits when you drag and drifts back to the cinematic rail when you let go, the way Meng To’s Sakura River Valley does. Its rain near the end of the loop was the other idea I borrowed.
  • Assets.
    • Poly Haven CC0 scans for rocks, cliffs, mud and roots.
    • Atlantic Forest species that have no scan (juçara, jerivá, embaúba, ipê, figueira, aninga, caeté), generated with Mint and rebuilt in Blender with LODs and impostors.
    • Field recordings, CC0 and CC-BY only. The crew stays silent.

The people fought hardest

If you’re asked which historical detail was hardest, the honest answer is the humans. The sources are precise, and humans are where the uncanny valley bites.

My first try was an AI 3D generator. It had two problems:

  • It refused. Any prompt that described the people accurately, and they wore very little, tripped its moderation.
  • Detail. What it did produce was fine for a canoe fifty metres away and poor in close-up: a mesh that reads as a figurine.

So the bodies come from MakeHuman, generated headlessly through the MPFB2 add-on in Blender by a script. Everything historical is built on top of them in code:

  • a skin shader;
  • procedural genipap and urucum paint masks;
  • a tonsure made of shell hair cards;
  • the enduape plume, cords, armbands and the tembetá;
  • a paddling rig.

An AI-generated face next to a MakeHuman face, and the MakeHuman face with genipap, urucum and tonsure

A MakeHuman warrior in lookdev with body paint, cords, feather ornaments and the tembetá

It’s still the least convincing part of the film. The crew reads as people at paddling distance and as well-painted mannequins up close.

The Gauntlet

A single agent produces one decent result and stops, because it grades its own work and knows every reason behind every choice. So I reused the loop from my browser FPS: the Gauntlet.

  • Builders. They run in parallel, each owning its own files.
  • Critics. They start with a clean context and only ever see the rendered pixels, never the builder’s explanation.
  • Classifier. A “does this read as real?” pass scores 42 frames of the film every round.
  • A/B captures. They verify every claimed fix.
  • Regression hunter. Its only job is to find what got worse.

I ran it overnight, unattended. The “reads as real” average went from 4.96 to 5.46 out of 10 over 13 rounds. That’s a small number, and an honest one. Most of the gains were artefacts that stopped showing up, not scores going up:

Round 4 with glare on the water and flat mist, next to the final frame

Three things the loop taught me:

  1. A builder saying “fixed” is not evidence. Several fixes were wrong, and only the A/B capture showed it.
  2. Capture bugs masquerade as scene bugs. For a while the crew looked see-through. It wasn’t a render bug: auto exposure, fog and contrast together were washing them out. Another time a “tiled frame” detector was reading an empty WebGPU canvas instead of the screenshot. Measure the measuring tool first.
  3. Root causes are rarely where they look.
    • The “low-poly rocks” were real scans silently dropped by a temporal-dead-zone bug. What was on screen were cliff meshes stretched vertically.
    • The “foam collars” on rocks were ground mist.
    • A yellow cross in the sky was a dragonfly an inch from the lens.

”It froze on the loading screen”

The first real-world report came from a friend on a modest home PC: the browser froze on the loading screen. I measured it:

  • Where the time went. The first frame built almost every shader in the scene in one task, 5.2 s of blocked main thread even on an M-series Mac.
  • Why it froze. On a Windows PC, where the driver compiles shaders much more slowly, that’s long enough to look like a hang.
  • Where the progress bar was. Already at 100 %.

The fixes, all measured before and after:

  • Shader warm-up. Objects are introduced a few at a time behind the loading screen, and every step waits for the GPU. No task runs long enough to freeze the browser.

  • Baking. The terrain noise, the crew’s adornment skinning and the animals’ placement searches were deterministic but ran in every visitor’s browser. They’re now computed once at build time and shipped as binary files.

    WhatAt load, beforeAfterBaked data
    Terrain2.8 s0.18 s6.8 MB
    Crew adornment skinning1.67 s0.39 s0.6 MB
    Animal placementa 3 s stall0.09 s0.1 MB

    Baked against computed at runtime: 0.00 mean pixel difference over 12 frames.

  • Device tiers.

    • Integrated GPUs, ≤ 4 GB of RAM or ≤ 4 cores, and phones get a lighter preset.
    • Software rasterisers get the recorded film instead of the live scene.
    • A boot that really hung steps the next visit down.
  • A cold open. The loading screen plays a short loop of the canoe behind the title, so there’s something to watch while the shaders compile.

One idea backfired, and I’m leaving it in because it’s the useful part. I tried lazy-loading the forest and mangrove after the film started. The boot got faster, but on the live site the film then stuttered for about 20 s while their shaders compiled mid-shot. Compiling them asynchronously fixed the stutter, but the trees then popped in 40 s late. The winning trade-off was the least clever one: the forest compiles up front, behind the loading screen, and only the animals, the ibises and the smoke, which first appear more than a minute in, load in the background.

End result on an M-series Mac: module loading went from 4.9 s to 1.5 s, and the film starts at about 7.7 s with no hitches after. It used to be 11–16 s followed by 20 s of stutter.

Crabs and herons

After the performance work shipped I got the next review: “in the crab close-up the crabs are totally low poly, and there’s a white bird with a black head you can’t identify.” Both were placeholders: a procedural crab with an eight-sided shell and stick legs, and a black-crowned night heron (savacu) built as a white ball with a black ball on top. The species were right, the models weren’t.

They were replaced with Mint generations, cut up in Blender:

  • A script splits each crab into carapace, claws and legs, so they patter on jointed legs.
  • The fiddler’s big claw actually waves. The females, which have two small claws, are new.
  • The night heron turns its head and shifts its weight.

Before and after: the procedural crab and the new uçá crab

Before and after: the procedural herons and the new cocoi heron and night heron

What it cost

ModelClaude Opus 5.5, 1 orchestrator + 31 subagents
Calls~3,770
Tokens~1.29 billion processed, mostly cache reads; ~0.7 million generated
API-equivalent price≈ US$ 516 at list price
Other AI~US$ 6 of Mint credits for the tree species
Wall clock~22 hours from first prompt to the first public build, including one unattended night

The performance and fauna work afterwards added five more agents. The cost is dominated by cache reads: 32 agents re-reading a large project many times. Token-wise, generating the code was the cheap part. Checking it was the expensive part, and it’s also what made the result trustworthy.

What’s still weak

  • The crew. Real-looking bodies with historical ornaments are the hardest problem here. A MetaHuman or photogrammetry pipeline is the next step if I go further.
  • Frame rate. It isn’t a solid 60 fps. On a loaded M-series Mac it’s around 40–45 fps, with dynamic resolution doing a lot of work.
  • Shader compilation. It’s still most of the boot. Merging materials that differ only in parameters into shared shaders is next.

Open Tupi on a desktop with WebGPU and turn the sound on. The prompt that built it is on the demos page.