Quen 3.8 Max: Multimodal AI with 1M Token Context & Low Cost

 4 min video

 2 min read

YouTube video ID: ppQh4Tc9BmM

Source: YouTube video by Two Minute PapersWatch original video

PDF

The AI landscape is rapidly evolving, with new open-source models challenging established players and offering unprecedented capabilities at significantly lower costs. DeepSeek Flash has emerged as a fast and affordable AI system, but the open frontier models, particularly Quen 3.8 Max, are making waves by directly competing with offerings from OpenAI and Anthropic.

Quen 3.8 Max: A Game Changer

Quen 3.8 Max stands out with several impressive features:

  • Multimodal Capabilities: It possesses "eyes and ears," indicating its ability to process various forms of data.
  • Extensive Context Window: A 1 million token context window allows it to handle vast amounts of information.
  • Agentic Workflows: This is a key highlight, as the model can operate independently for extended periods. Demonstrations show it working autonomously, writing, testing, and repairing its own code, and even reproducing and improving research papers. One notable example involved it thinking independently for 16 days, starting from an empty folder.
  • Cost-Effective: Its API pricing is significantly lower, potentially five to ten times cheaper than competitors, which could force other AI providers to reduce their prices.
  • Website and App Creation: It can effortlessly create impressive websites and applications.
  • Commitment to Openness: The developers have committed to releasing the model's weights soon, fostering further innovation in the open-source community.

While the full Quen 3.8 Max model is gigantic and may not be runnable on personal hardware, the good news is that a variety of smaller models are also available.

The "Toyota Corolla" of AI: Smaller Quen Models

Earlier, smaller Quen models like 3.6, 27, and 35 billion parameters have achieved legendary status. Despite being several months old, many consider them the best in their categories, akin to the "Toyota Corolla of the AI world." These models are accessible to individuals with more modest resources and could become the next "daily driver" for many, available for free.

Humanity's Last Exam: A Benchmark for Real-World Performance

When evaluating AI models, benchmarks are crucial. "Humanity's Last Exam" is highlighted as a particularly difficult and indicative academic benchmark for real-life performance. Initially, the best billion-dollar closed AI systems could only achieve about 2% on this exam. However, just over a year later, an open model has surpassed 50%, demonstrating the rapid progress in open-source AI. This benchmark is recommended for tracking the true capabilities of AI systems.

The Golden Age of Open Science and AI

The current era is described as a "golden age of open science and open AI systems," with incredible new gifts emerging weekly. Quen's contributions to open science and open source are particularly lauded.

For those looking to experiment with AI, platforms like Lambda offer powerful NVIDIA GPUs to run experiments, train models, fine-tune existing ones, perform inference, or generate text-to-image/video content. This allows researchers and enthusiasts to quickly test ideas and reproduce AI research papers.

  Takeaways

  • Quen 3.8 Max offers multimodal processing (“eyes and ears”) and a massive 1 million‑token context window, enabling it to handle extensive data streams in a single prompt.
  • Its agentic workflow capability lets the model operate autonomously for days, writing, testing, and repairing code and even improving research papers without human intervention.
  • The API pricing is projected to be five to ten times cheaper than comparable OpenAI or Anthropic services, potentially forcing the industry to lower prices.
  • Smaller Quen models (3.6‑35 B parameters) are free and run on modest hardware, earning the nickname “Toyota Corolla of AI” as reliable daily drivers for many users.
  • The “Humanity’s Last Exam” benchmark shows open‑source models jumping from ~2 % to over 50 % accuracy within a year, highlighting the rapid progress of open AI in real‑world tasks.

Frequently Asked Questions

What does “agentic workflows” mean in the context of Quen 3.8 Max?

Agentic workflows refer to the model’s ability to act independently over extended periods, initiating and completing tasks without continuous human prompts. In demonstrations, Quen 3.8 Max autonomously wrote, tested, and repaired its own code and even refined research papers, showcasing self‑directed problem solving.

How does the 1 million token context window improve Quen 3.8 Max’s capabilities?

A 1 million token context window lets Quen 3.8 Max retain and process far larger text blocks than typical models limited to a few thousand tokens. This enables it to analyze extensive documents, maintain long‑term reasoning, and generate coherent outputs across massive inputs, reducing the need for chunking or external memory management.

Who is Two Minute Papers on YouTube?

Two Minute Papers is a YouTube channel that publishes videos on a range of topics. Browse more summaries from this channel below.

Does this page include the full transcript of the video?

Yes, the full transcript for this video is available on this page. Click 'Show transcript' in the sidebar to read it.

Helpful resources related to this video

If you want to practice or explore the concepts discussed in the video, these commonly used tools may help.

Links may be affiliate links. We only include resources that are genuinely relevant to the topic.

Full transcript is not shown on this page

This page focuses on the summary and original notes. For full verification, refer to the original YouTube video.

PDF