Ox Alpha (GLM 5.3 Flash) Overview: Features, Pricing, and Use Cases

 5 min video

 4 min read

YouTube video ID: r-tzcMlQISk

Source: YouTube video by FireshipWatch original video

PDF

A new frontier-grade model, dubbed Ox Alpha, recently appeared anonymously online and quickly became the most popular model on Open Router, serving 42 trillion tokens in its first six days. This model boasts a million-token context window, accepts video input, and excels at writing code. Initially, it was completely free, leading to speculation about its origin, with some guessing it was from Jepu, others Xiai, and some even suggesting Google.

The Anonymous Model's Origins

The model's anonymous launch sparked significant interest, especially given the recent impact of Chinese AI developments like DeepSeek Harness. Several clues pointed to a Chinese origin for Ox Alpha:

  • Precedent: Ox Alpha was the fifth anonymous "animal model" to appear on Open Router, and the previous four were all confirmed to be from China.
  • Technical Footprints: Developers noticed that its stack traces and error codes matched Zoo's API, and a tokenizer fingerprint test aligned with Zoo's GLM series.

Despite the fine print indicating that all prompts would be retained by a "mysterious, probably foreign entity," users readily adopted the free and highly effective model. Developers reportedly pasted proprietary company code into it and used it to bypass rate limits on other services like Claude. At its peak, Ox Alpha accounted for nearly a third of Open Router's weekly traffic.

The hype surrounding Ox Alpha led to some misinformation, including a viral but misleading screenshot showing an 80% score on the DeepSeek SWE benchmark, when the actual score was closer to 58%. There was even speculation that it might be a stolen Claude checkpoint.

The Reveal: GLM 5.3 Flash

On August 26th, Jepu officially announced that Ox Alpha was indeed GLM 5.3 Flash. This model is a natively multimodal mixture-of-experts model with 320 billion total parameters. Its MIT-licensed weights were released on Hugging Face the same day. A significant detail that surprised many in Silicon Valley was that the entire stealth operation was run on just 100,000 Chinese-made chips.

The reveal also marked the end of the free access period, with the anonymous endpoint replaced by a pricing page. However, the pricing remains highly competitive: 15 cents per million input tokens and 50 cents per million output. With a 50% discount available until September 9th, it is currently up to 40 times cheaper than Claude. Unlike many new model launches that come with unverified benchmarks, GLM 5.3 Flash's benchmarks have held up well under external scrutiny.

Practical Application: Modernizing "Horse"

Despite its impressive capabilities, the model has some drawbacks: it can be slow, verbose, and occasionally gets stuck in "doom loops." To truly test its effectiveness, it was applied to a personal project: "Horse," a 2016 Angular.JS app with its last commit a decade ago. The goal was to modernize the legacy stack while preserving its original humor and "vibes."

The model successfully detected the outdated technologies and meticulously narrated its plan of attack. It then proceeded to build a design system using raw CSS, featuring a purple slot background and rainbow gradient buttons, reflecting a unique aesthetic. Its vision capability proved valuable by detecting and fixing a CSS overflow bug on mobile screens using the "min-width equals zero" trick. It also added subtle touches like enter animations for video cards and new humorous elements, such as toast notifications.

Code-wise, it replaced the legacy stack with a "no stack" approach, writing the app in HTML, CSS, and vanilla JavaScript, resulting in code that ironically resembled a new legacy codebase.

Vision Skills Test: Video Analysis

To further test its vision skills, the model was given a random horse video and asked to integrate it into the app, creating a relevant title, description, and comments. The model utilized FFmpeg to break down the video into frames, analyzed them, and accurately grasped the full context of the video. It generated a fitting description and aligned the humor with the platform's existing tone.

Exa: A Search Engine for AI Agents

The article also highlights Exa, a search engine designed for AI agents. Exa offers more than just web results, providing access to 350 million research papers, financial data, and information on over 1 billion people and 70 million companies, with new sources added weekly.

Exa was integrated into a project called Polyhip to allow users to gamble on the stock market, including betting on tech CEOs selling company stock before earnings calls. Exa successfully pulled recent SEC filings and converted them into structured JSON. It also cites all its sources, enhancing transparency. Exa can be directly integrated with Claude, ChatGPT, or other AI agents via native plugins. The Exa Agent feature allows for more complex queries that require data from multiple datasets simultaneously, offering greater power than standalone LLMs. Exa offers a free trial with $20 in credits.

  Takeaways

  • Ox Alpha, later identified as GLM 5.3 Flash, quickly became the top model on Open Router, processing 42 trillion tokens in six days with a million‑token context window and video input capability.
  • The model’s anonymous launch and technical fingerprints—stack traces matching Zoo’s API and a tokenizer aligned with Zoo’s GLM series—pointed to a Chinese origin, later confirmed by Jepu’s announcement.
  • GLM 5.3 Flash is a 320‑billion‑parameter multimodal mixture‑of‑experts model released under an MIT license, priced at $0.15 per million input tokens and $0.50 per million output tokens, with a 50 % discount until September 9.
  • In a practical test, the model modernized a legacy 2016 AngularJS app called “Horse” by detecting outdated tech, generating a new design system, fixing CSS bugs, and rewriting the codebase in plain HTML, CSS, and vanilla JavaScript.
  • The article also showcases Exa, an AI‑agent‑focused search engine that provides access to billions of data points and can be integrated via plugins, demonstrated by pulling SEC filings for a stock‑betting project.

Frequently Asked Questions

What evidence linked Ox Alpha to a Chinese origin before its official reveal?

Analysts noted several clues that pointed to China, including the fact that Ox Alpha was the fifth anonymous “animal model” after four confirmed Chinese releases, its stack traces and error codes matched Zoo’s API, and a tokenizer fingerprint test aligned with Zoo’s GLM series, all suggesting a Chinese developer.

How does GLM 5.3 Flash’s pricing compare to Claude’s cost?

GLM 5.3 Flash costs $0.15 per million input tokens and $0.50 per million output tokens, and with a 50 % discount through September 9 the effective rate is up to forty times cheaper than Claude’s pricing, making it one of the most affordable high‑performance multimodal models on the market.

Who is Fireship on YouTube?

Fireship is a YouTube channel that publishes videos on a range of topics. Browse more summaries from this channel below.

Does this page include the full transcript of the video?

Yes, the full transcript for this video is available on this page. Click 'Show transcript' in the sidebar to read it.

Helpful resources related to this video

If you want to practice or explore the concepts discussed in the video, these commonly used tools may help.

Links may be affiliate links. We only include resources that are genuinely relevant to the topic.

Full transcript is not shown on this page

This page focuses on the summary and original notes. For full verification, refer to the original YouTube video.

PDF