TechNewsReel
Live

Harness Choice Drives Success Rates in Three.js AI Coding Benchmark

Research by alvins82 reveals that AI harnesses significantly impact the success rates and token efficiency of LLMs during complex 3D web development tasks.

TechNewsReel Newsroom · September 8, 2026

A new benchmark conducted by researcher alvins82 highlights the critical role that 'harnesses' play in the performance of large language models (LLMs) when executing complex 3D coding tasks. The study analyzed the ability of various AI combinations to generate a sophisticated, self-contained Three.js environment, revealing a wide variance in success rates and resource consumption.

To test these capabilities, the researcher used a single, detailed prompt: "Build a single-page Three.js sci-fi hangar with hovering drones, animated warning lights, emissive runway strips, and subtle volumetric-style fog planes. Include drone formation toggle and cinematic camera path." The tests were performed in '/goal' mode across ten different combinations of models and harnesses. The models tested included GLM 5.3 Flash Max, Luna 5.6 Max, SOL 5.6 Max, Astra 6.0 Max, and Qwen 3.8 27B x-high, paired with harnesses such as Codex, OMP, OpenCode, and DSH/PTC.

The Performance Gap

The results indicate that the choice of harness can be as influential as the model itself. For instance, when using the GLM 5.3 Flash Max model, the 'OpenCode' harness achieved a high success and completion rate of 96.89%. In contrast, the same model paired with the 'Codex' harness saw the lowest performance in the group, with a success rate of 85.26%.

Beyond success rates, the researcher tracked several key performance metrics to determine efficiency, including total duration, active turn time, input and output tokens, reasoning tokens, and the number of tool errors encountered. The data revealed a massive disparity in token usage. The GLM 5.3 Flash Max model using the OpenCode harness recorded the highest consumption with 4,355,915 total tokens, while Qwen 3.8 27B x-high using the OMP harness followed with 3,479,386 total tokens.

Industry Implications

These findings suggest that for specialized tasks like Three.js development—which requires a precise blend of mathematical logic and creative asset placement—the infrastructure surrounding the model is a primary driver of reliability. The fact that the same model can swing from a 96% success rate to 85% based on the harness suggests that prompt orchestration and tool-calling frameworks are currently a bottleneck in AI-driven software engineering.

Future Outlook

As developers increasingly rely on LLMs for boilerplate and complex 3D scene generation, optimizing these harness-model pairings will be essential to reduce token costs and error rates. Future research may examine whether these performance gaps persist across different coding languages or if the 'OpenCode' harness provides a systemic advantage for all 3D-related tasks. For now, the data underscores that selecting the right model is only half the battle; the delivery mechanism is equally vital.

Get a notification when a big story breaks. A few a day at most — no spam.