Full-StackPrototype2026
Wildwood
A voice-controlled first-person survival game in the browser. Say "make it night and grab my flashlight" and TypeSafe's Jev turns the sentence into game actions in about a tenth of a second.
- 115 ms
- Median Jev time per command, measured
- 19
- Typed questions per sentence
- $0.06
- Cost per thousand commands
Timeline
Oct 2026
Role
Game design, 3D world and voice pipeline
Team
Solo
Domain
Gaming · Voice interfaces
Status
Prototype
Stack
Source
Section 01
Overview
Wildwood is a survival sandbox in a forest valley: walk with the keyboard and mouse, and change the world by talking to it. Weather, time of day, what is in your hands, using supplies, lighting the campfire and reloading all happen by voice, with no inventory menus or weapon wheels.
The idea behind it: mapping a loose sentence onto a fixed set of game actions is a classification problem, so a fast decision model fits better than an LLM writing a tool call.
Section 02
Voice to Action
The browser's speech recognition turns speech into text. Each finished phrase goes to the server with the current game state, and one Jev request asks 19 typed questions: six yes/no gates (does the player want to change the weather, the time, what they hold, use an item, act on the world, or check status?) and a pick for each category.
Plain TypeScript then applies the rules: a category fires above 0.5, only one healing item is used per need, time always moves forward, and two actions in one sentence complete the combo objective. The HUD shows the parsed actions and how long Jev took.
Section 03
The World
The valley is rendered with Three.js through React Three Fiber: up to 190,000 GPU grass blades, thousands of distant impostor trees, a cabin, a campfire, day-night lighting, fog, rain, thunderstorms and snow that eases in over a few seconds. Models and textures stream from Poly Haven, and three power modes trade detail for frame rate.
Section 04
Measured Results
Over 28 commands, the Jev request took a median 115 ms (73 to 170 ms). From pressing Enter on a typed command to the HUD updating took a median 190 ms. Every request used about 1,330 input tokens, roughly $0.000056 per command. Spoken commands add the browser's own speech recognition time on top of these figures.
The How Jev Hears You page breaks a sentence into gates and picks and compares the frame budget with an LLM tool call; that LLM figure is a labelled estimate.






