Claude Opus 5 Is a Freak: The Ultimate Guide to Its Wild AI Coding, 3D and Agentic Workflow Results

Futuristic illustration of an AI core building a browser-style operating system workflow with split visuals showing smooth capability on one side and slow, glitchy performance on the other, including subtle 3D modeling elements.

Claude Opus 5 is one of the most capable and one of the most frustrating frontier AI models tested so far. It can build a functional Windows 11-style operating system inside a browser from a single prompt. It can create a 3D office scene from an image, control Blender to build an X-Wing fighter, research corporate earnings, generate a narrated financial presentation, and operate a DAW to compose a full electronic track.

That is the good news. The bad news is that it is painfully slow, expensive, and not consistently ahead of competing models in the areas that matter most to many businesses. For Canadian companies evaluating AI tools, that distinction matters. A model that produces an impressive demo is not automatically the model that delivers the best return on investment for everyday software, marketing, research, or operations work.

The real story is not that Opus 5 is universally superior. It is that agentic AI workflows are becoming absurdly powerful. AI systems can now plan tasks, access local files, use software tools, verify outputs, identify bugs, and revise their own work. That changes the equation for teams across the GTA, Canadian startups, enterprise IT departments, creative studios, and organizations trying to stay ahead of the AI curve.

Opus 5 intro

Claude Opus 5 is Anthropic’s latest high-end frontier model, built not just for chat but for agentic workflows. That means it is designed to tackle long-horizon work involving multiple steps, software tools, file systems, web searches, coding environments, and iterative testing.

For simple questions, Claude.ai may be enough. But the more revealing test is using Opus 5 through an agentic harness such as Claude Code. This gives the model access to multiple local files and folders, enabling it to work across projects instead of merely returning text in a chat window.

This is a major distinction for Canadian business technology teams. The value proposition is not, “Can this AI write an email?” It is, “Can this AI execute a multi-stage workflow that would normally require a developer, designer, analyst, researcher, and QA tester?”

Opus 5 clearly has the raw capability to do that in certain categories. It plans extensively, creates structured project files, runs local applications, checks its own work, and tries to repair issues without needing repeated human intervention. That is seriously impressive.

But capabilities need to be separated from operational practicality. Several of the major tasks took close to, or more than, an hour. In production environments, speed, price, governance, and reliability are every bit as important as a jaw-dropping one-shot result.

Windows on browser

The first major test was a monster prompt: create a browser-friendly recreation of Windows 11, complete with common applications, a taskbar, Start menu, File Explorer, Microsoft Office-style tools, Microsoft Store-style downloads, media apps, Discord, Slack, Spotify, Paint, Sticky Notes, a code editor, and games.

The request also required those programs to actually function while running efficiently in a standard browser.

Opus 5 did not simply produce a static mock-up. It created a working browser-based environment with core desktop behaviour, including:

  • Desktop, taskbar, Start menu, and login screen interactions
  • Window focus, resizing, dragging, snapping, and layout mechanics
  • Dark mode, light mode, brightness, night light, and calendar controls
  • A Word-style editor with formatting, text alignment, font colour, and saving
  • An Excel-style spreadsheet with working average and sum formulas
  • A file system that retained saved documents and spreadsheets
  • A simulated Microsoft Store that could “download” additional apps
  • Functional browser versions of Paint, Sticky Notes, Solitaire, and File Explorer

The Word-style document could be edited, formatted, saved with Ctrl+S, closed, and reopened with the changes intact. The spreadsheet tool could calculate formulas and save a new file that appeared later in the Start menu. That is a long way beyond basic vibe coding.

Some applications were more convincing than others. PowerPoint-style functionality was limited. Text boxes could not be moved freely, and many of the capabilities expected from real presentation software were missing. Spotify-style music was basic, and Discord and Slack responses were simulated rather than genuinely context-aware. The interface looked right, but the conversational behaviour was largely hard-coded.

Still, the scale of the result is wild. Opus 5 planned the application architecture, built the components, launched the browser page, detected bugs, and attempted fixes autonomously. The project consumed roughly 366,000 tokens and ran for more than an hour.

For Canadian software leaders, the lesson is clear: this class of AI can dramatically accelerate prototyping. A product team in Toronto or Vancouver could use a workflow like this to validate interfaces, internal tools, training simulations, or proof-of-concept software before investing in a conventional build. But nobody should confuse a compelling prototype with a secure, scalable, production-grade operating system.

Image to 3D

Next came a difficult image-to-3D challenge. Opus 5 received a reference image of a geometric office scene and was asked to create a faithful, animated 3D version in a single HTML file.

The initial output was decent but too basic. Object placement was off, the glow was excessive, and many key details were missing. After a single corrective instruction requesting closer fidelity to the image, the model produced a much stronger result.

The revised scene captured most of the desk and chair arrangement, recognized some people in the space, and generated a layout that was visibly closer to the reference than comparable outputs from another frontier model. It was not perfect. There were extra chairs, inconsistent books, misplaced wall panels, and missing desks.

Yet it was still the strongest result in this comparison. Opus 5 showed a notable strength in spatial composition, frontend rendering, and iterative visual refinement.

This task consumed approximately 343,000 tokens and again approached an hour of runtime. For architecture, commercial real estate, retail design, workplace planning, marketing, and product visualization, the opportunity is obvious. Canadian firms could generate early-stage visual concepts quickly. The limitation is equally obvious: precise design work still needs human review because visual inaccuracies can matter a lot when a project involves brand standards, safety, construction, or client approval.

Financial presentation video

The financial presentation test pushed Opus 5 into a more business-relevant direction. The prompt asked it to find the previous year’s Q4 reports for Nvidia, Google, Meta, and Amazon, analyse the financials and outlooks, then create a roughly one-minute 16:9 presentation video with charts, visuals, background music, motion graphics, and a Gemini TTS voiceover.

No earnings reports were attached. Opus 5 had to locate the relevant materials, process the figures, install the Hyperframes motion graphics tool, generate narration through Gemini TTS, render the final video, and verify the output.

The finished presentation made a concise argument around three themes:

  • Scale: Amazon generated the largest revenue figure, at approximately US$213 billion in the comparison.
  • Growth: Nvidia posted the fastest growth rate, cited at 73%.
  • Profitability: Meta stood out for operating profitability, while Nvidia’s gross margin was highlighted at 75%.

The video also framed AI infrastructure as the common engine beneath all four companies: Nvidia’s data centre business, Google Cloud, AWS, and Meta’s enormous AI investment cycle. It identified the enormous capital spending expected from Alphabet, Meta, and Amazon, then contrasted that spending with Nvidia’s position as a major infrastructure supplier.

The core question was sharp: the spending is happening, but who earns the return?

For CFOs, CIOs, analysts, and strategy teams, this is the type of agentic workflow that should get attention immediately. The model did not just summarize an earnings release. It assembled research, visual storytelling, narration, and production. That could be useful for internal briefings, sales enablement, competitive intelligence, investor communications, and executive updates.

However, financial analysis must be verified. Any organization using AI-generated content around financial performance, forecasts, securities, or investment decisions needs human validation of source documents, calculations, dates, and conclusions. A clean video is not a substitute for analytical accountability.

Luma Agents

If the same creative workflow is needed repeatedly, Luma Agents offers a more repeatable framework. The central concept is straightforward: instead of rebuilding an AI workflow every time, turn a successful process into a reusable Luma Skill.

For example, an image can be used as a reference to generate a complete brand kit and colour palette. Once the output process works well, the workflow can be saved as a skill called something like “Brand Kit.” Later, another image can be uploaded and the same workflow can be reused.

A fashion example makes the point even clearer. An uploaded photo of a person and an outfit can be turned into a fashion-show-style video. Save that workflow as a “Fashion Show” skill, then use a new model or new clothing item to repeat the process without rebuilding the pipeline from scratch.

For Canadian marketing teams and creative agencies, this is potentially valuable because repeatability is where AI begins moving from novelty to operations. Brand production requires consistency. A team cannot have every campaign generated through a completely different ad hoc process.

Luma Skills also support collaboration. A workflow can be shared by link, installed by teammates, or bundled with related skills into a larger creative toolkit. That could help organizations standardize campaign assets, product visualizations, social content, video formats, and brand workflows across distributed teams.

The important idea is bigger than any individual platform. The future of business AI is not just prompting. It is packaging reliable methods into reusable systems.

Blender X wing fighter

Opus 5 was also connected to Blender through a local MCP setup and asked to create an X-Wing fighter with realistic textures and motion.

The model built the craft step by step. It generated the body, wings, components, and hinges needed to make the wings open. It also created an animation and used screenshots to inspect whether the model and motion behaved correctly. If it found problems, it could move into a repair cycle.

The final model included detailed wing assemblies, materials, textures, and even a small compartment for R2-D2. For a generated asset created from scratch through tool use, it was extremely comprehensive.

This task used about 173,000 tokens, much less than some of the others, though it still took close to an hour. The model’s ability to coordinate with specialist software is one of its most important features. It is no longer limited to describing a 3D asset. It can interact with the production environment used to create one.

That does not eliminate the need for 3D artists, technical directors, or quality control. It does mean those teams may soon spend less time on repetitive setup work and more time refining creative direction. In a competitive Canadian technology market, that could make smaller studios and product teams considerably more productive.

Music composition

The music test was another ambitious one-shot workflow. Opus 5 was asked to compose an award-worthy song in Waveform, a free digital audio workstation. It had to decide which instruments were needed, locate free VST plugins under 800 MB, download them, load them into the DAW, compose each track, configure effects, create panning and automation, then render a finished song.

The model first searched the computer for the DAW. It then researched plugin options. One plugin required account registration, so it could not be used. The model instead selected alternatives, including Surge XT, and asked for a musical direction. The chosen style was cinematic melodic techno.

The resulting arrangement was surprisingly detailed, with tracks for:

  • Kick, clap, hats, percussion, cymbals, and other drum elements
  • Sub bass and bass
  • Pad, air, strings, keys, arpeggiator, pluck, and lead sounds
  • Countermelody, risers, and effects automation
  • Reverb, delay, panning, and additional production details

The finished song was five minutes long and reasonably polished, though musically repetitive, with similar chords carrying much of the track. Still, Opus 5 successfully navigated the DAW, installed instruments, arranged MIDI content, set automation, and exported a listenable piece of music from one prompt.

That is nuts. It is also another example of why the conversation around AI needs to evolve. The question is not whether a model can create a rough loop. The question is whether it can coordinate a complete sequence of tool-based creative decisions.

The trade-off again was speed. The workflow consumed roughly 350,000 tokens and took over an hour. For a business that produces high volumes of generic audio, internal demos, quick prototypes, or temporary media assets, this may become useful. For high-stakes commercial music, the human creative layer remains essential.

FROG TEST

The frog test is a deceptively simple visual reasoning challenge. An image contained a hidden frog, but the prompt only asked whether there were any animals in the image and requested that they be identified and circled.

Opus 5 tried hard. It split the image into a three-by-three grid, examined each area, considered whether it might be seeing a snake, adjusted saturation and contrast, and rescanned individual tiles. It ultimately failed to identify the frog.

That is not great. But there was one positive detail: the model did not confidently invent an animal it could not find. Instead, it acknowledged the uncertainty.

For enterprise AI use, that is a meaningful distinction. A model that fails transparently is safer than one that produces a polished but fabricated answer. Still, a failure is a failure. Businesses should not assume that impressive coding and design performance automatically translates into dependable visual inspection capability.

Identifying cancer

The medical imaging test produced an even more serious warning. Opus 5 was given six scans containing tumours and asked to identify the tumour types. It got all six wrong.

Its classifications included incorrect diagnoses, missed tumours, uncertainty around the wrong possibilities, and statements that no discrete mass was visible where a tumour was present. Another frontier model got one image correct in the same comparison, but that is hardly a reassuring standard.

Opus 5 was more willing than some systems to engage with the biomedical prompt. However, willingness to answer is not clinical accuracy. This test reinforces a fundamental principle for Canadian healthcare organizations, medtech firms, research institutions, and any business handling medical data: generative AI must not be treated as an autonomous diagnostic authority.

Medical imaging interpretation requires qualified professionals, validated tools, appropriate regulatory processes, and rigorous clinical evidence. A model that can make a beautiful website or compose music can still be catastrophically unreliable in medicine.

Deep research

Opus 5 performed better when asked to conduct deep biomedical research on the pathophysiology of atherosclerosis. It answered the request, generated structured explanations, created a flowchart, coded diagrams, presented comparison tables, and included a bar graph.

The output was useful, but not the preferred result in this comparison. Other models, including GPT 5.6 and Kimi K3, were considered more thorough while remaining concise. Gemini tended to be more verbose. Opus 5 felt somewhat restrained, with less complete organization and depth than the strongest alternatives.

This is subjective, but it matters. Business users rarely need the “best model” in the abstract. They need the model that gives them the most useful output for a specific task. For research-heavy Canadian organizations, that could mean comparing models by domain, output style, cost, and the quality of citations or source handling, not simply adopting the highest-profile option.

Specs cost benchmarks

Opus 5 is built for long, multi-step work and supports a one-million-token context window. That is roughly 700,000 words or a small-to-medium-sized codebase. For large technical documentation sets, extensive project files, or complex agentic processes, that context capacity is substantial.

Anthropic’s reported benchmarks suggest Opus 5 improves on its prior top model across many categories, though the earlier model reportedly performs better in legal and health tasks. Opus 5 performs better in biology.

Benchmark results, however, are not a clean victory lap. On DeepSuite 1.1, a benchmark intended to measure agentic software engineering, GPT 5.6 reportedly achieved the highest score in the cited comparison. On Artificial Analysis, Opus 5 ranked first, but only marginally ahead of several competitors. Without confidence intervals, it is difficult to know whether those small gaps represent meaningful performance differences.

Opus 5 also reportedly achieved more than 30% on ARC-AGI-3, a benchmark involving learning rules inside game environments. That is far ahead of the next cited model. But there are concerns that the model may be optimized for benchmark-like tasks. When tested with unfamiliar games containing unusual rules, it reportedly performed poorly, even below Opus 4.8 in some cases.

The most practical findings are simpler:

  • Speed: Opus 5 is extremely slow, often slower than GPT 5.6, Kimi K3, and even the previous Fable 5 model.
  • Cost: It is significantly more expensive, approaching twice the cost of GPT 5.6 Sol in the cited comparison.
  • Hallucinations: Its hallucination rate is roughly similar to Kimi K3, while GLM 5.2 was cited as hallucinating substantially less.
  • Livebench: Opus 5 ranked third, behind GPT 5.6, across reasoning, coding, mathematics, data analysis, language, and instruction following.
  • Debate: It ranked first on the cited adversarial multi-turn debate benchmark.
  • Finance and coding: In the cited VALS index, it trailed Fable 5 and only narrowly exceeded Kimi K3 while costing much more.

The bottom line is brutal: Opus 5 is excellent, but it is not clearly excellent enough to justify its premium for every organization. Canadian technology leaders should calculate the value per completed workflow, not just the price per token or the position on a leaderboard.

Guardrails

Like other Anthropic models, Opus 5 has guardrails around cybersecurity and biology. It may reject certain prompts or fall back to a less capable model, Opus 4.8, when a request falls into restricted territory.

It is described as less restrictive than Fable 5. It can help find vulnerabilities in source code, for example. It is also more permissive on some biology-related questions, as shown by its willingness to answer the biomedical research request.

However, there are still restrictions around potentially risky long-running autonomous biology research and other sensitive areas. For Canadian enterprises, this creates a procurement and governance issue. AI teams need to know not only what a model can do, but also when it may refuse a request, silently route work to a different model, or produce a materially different result than expected.

Opus 5 requires a paid Claude plan for access through the Claude interface and Claude Code. It is also available through an API.

Verdict

Claude Opus 5 is a freak in the best sense of the word. Its frontend development, visual design, 3D generation, software tool use, and autonomous workflow execution are seriously impressive. It produced some of the cleanest and lowest-error creative outputs in these tests.

But it is also very slow and very expensive. In many normal workflows, GPT 5.6, Kimi K3, or even a lower-cost option such as GLM 5.2 may be sufficient. The performance gap often looks small, while the cost gap can be enormous.

That makes Opus 5 a specialist tool rather than an automatic default. It is most compelling when the task involves:

  • High-end frontend development and interactive prototypes
  • Complex browser applications and vibe coding
  • 3D scenes, Blender automation, and visual generation
  • Long-horizon agentic tasks requiring planning and verification
  • Hard coding problems that competing frontier models cannot solve

For Canadian businesses, the smart move is not blindly standardizing on the most expensive frontier model. Build a model portfolio. Use the right system for the right task. Reserve premium models for high-leverage assignments where their strengths can create measurable value.

The AI race is accelerating at an insane pace. The organizations that win will not be the ones that chase every headline. They will be the ones that identify repeatable workflows, apply strong oversight, measure outcomes, and turn this new generation of AI agents into an actual competitive advantage.

Is your organization prepared to treat AI agents as part of its operating model, or are they still being used as glorified chatbots?

FAQ

What is Claude Opus 5 best suited for?

Claude Opus 5 appears strongest in agentic coding, frontend development, interactive browser applications, 3D design, Blender-based workflows, and complex multi-step tasks involving tools, files, testing, and revisions.

Is Claude Opus 5 faster than other frontier AI models?

No. In these tests, Opus 5 was notably slow. Many tasks ran for close to or more than an hour, and it was described as slower than GPT 5.6 and Kimi K3.

Can Claude Opus 5 create complete software projects from one prompt?

It can generate substantial prototypes from a single prompt. In one test, it created a functional Windows 11-style browser environment with desktop controls, office-style tools, saved files, apps, games, and automated bug-fixing. Production deployment would still require human engineering, security review, and testing.

Is Claude Opus 5 reliable for medical diagnosis?

No. In the tumour identification test, it incorrectly classified all six medical images. It should not be used as an autonomous medical diagnostic system.

Is Claude Opus 5 worth the cost for Canadian businesses?

It may be worth the cost for high-value frontend, 3D, or difficult agentic coding workflows. For many routine business tasks, less expensive models may provide comparable value at better speed and lower cost.

Does Claude Opus 5 have content restrictions?

Yes. It can restrict or reject some cybersecurity and biology-related requests and may fall back to a less capable model for certain sensitive tasks. Organizations should test important workflows before relying on them operationally.

Leave a Reply

Your email address will not be published. Required fields are marked *

Most Read

Subscribe To Our Magazine

Download Our Magazine