Open Model Resurgence and Olama’s Role in Enterprise AI Adoption

 57 min video

 9 min read

YouTube video ID: rY0wnfFHYbs

Source: YouTube video by Y CombinatorWatch original video

PDF

The landscape of AI is undergoing a significant shift towards open models, particularly within enterprise settings. This movement is driven by a combination of factors, including cost reduction, greater control over AI customization, and the rapid advancement of open-source models. Olama, a platform enabling the local and cloud-based execution of open-source AI models, stands at the forefront of this trend, boasting 9 million developers, 178,000 GitHub stars, and adoption by 85% of the Fortune 500.

The Driving Force: Cost and Customization

While cost is a primary motivator for enterprises adopting open models, it's not the sole factor. Businesses are increasingly seeking better control over AI and the ability to customize models for their specific needs. Solving the cost problem in the short term then enables them to pursue this long-term vision of tailored AI solutions.

AT&T, for instance, has already shifted 40% of its token consumption to open models, initially favoring US and European models but now evaluating Chinese alternatives. This demonstrates a clear trend towards leveraging open models for significant operational benefits.

The Rise of Coding Agents and AI Assistants

A major driver of open model adoption is the proliferation of coding agents and AI assistants. These tools, exemplified by projects like OpenClaw and Hermes, have dramatically increased token usage per developer and user. OpenClaw, in particular, has enabled non-developers in fields like finance, support, marketing, and sales to automate complex tasks, leading to an explosive growth in token consumption. This surge is also fueled by the expansion of context windows in open models, moving from 128K to over a million tokens.

Olama's cloud data reveals a 150x growth in token usage since the beginning of the year, largely attributed to the demand for open models. While custom fine-tuning of models like Deepseek and Kimmy was prevalent in 2023, the current trend shows a significant uptake of out-of-the-box open models.

The Cycle of Fine-Tuning and Model Releases

The interest in fine-tuning custom models has seen cycles. Early 2024 saw a peak, followed by a period of skepticism due to the rapid release of new models potentially rendering fine-tuned efforts obsolete. However, the trend is now returning, with the cadence of open-source model releases accelerating. This makes custom training more challenging but also highlights the improving tooling available for fine-tuning.

The potential slowdown in frontier model development due to AI safety concerns could further boost open-source models, which continue to advance rapidly.

Security and Safety: A Key Blocker and Opportunity

Security and safety remain significant concerns for enterprises considering open model adoption. However, if these issues can be addressed, the use of Chinese-origin models becomes a viable and exciting option for businesses. Interestingly, open-weight models have even been used by platforms like Hugging Face to detect hacks from frontier models.

One key advantage of open models lies in security testing. Unlike closed models that may refuse to perform penetration testing, open models can be specifically trained for such tasks, offering a powerful tool for ensuring software security. While some open models are custom-trained for more "liberal" security testing, even out-of-the-box models, with proper safety training, can be effective for legitimate use cases.

Olama's Role in the Ecosystem: Orchestrating Model Launches

Olama plays a crucial role in the open model ecosystem by coordinating model launches. With the increasing speed of model releases, Olama has developed a playbook for successful "day zero" launches, which involves:

  1. Inference Engine Support: Ensuring the model is supported by the inference engine, optimizing for speed and accuracy.
  2. Use Cases and Harnesses: Identifying appropriate use cases and preparing harnesses (like CodeX or OpenCode) to effectively run the new model and leverage its unique capabilities.
  3. Hardware and Providers: Collaborating with inference providers and hardware manufacturers (like Nvidia and Apple) to ensure optimal performance.

This process is often a "fire drill," with much of the work coming together in the 24 hours before a model's release. Olama acts as a crucial "operating system," integrating drivers, hardware, and application runtimes to provide a consistent and efficient experience for developers.

The Future of AI: Unbundling and Hybrid Models

The AI stack, often described as a "five-layer cake," has hidden layers that are ripe for unbundling into specialized companies. These include:

  • Knowledge: Connecting company data and context to models.
  • Coordination: Orchestrating sub-agents, some local, some in the cloud.
  • Execution: Solving compute problems for agents moving to the cloud.

This unbundling mirrors the evolution of cloud computing, where developers favored best-of-breed products over bundled solutions. While frontier labs aim for walled gardens, open-source developers are pushing for open alternatives, leading to a hybrid future.

The abundance of open model tokens is shifting the scarcity to what's above the tokens – how to orchestrate agents and solve complex engineering problems. The increasing capability of coding agents also suggests a future where software development becomes more accessible, potentially reducing vendor lock-in.

However, certain aspects, like managing stateful problems (e.g., storage, credentials, safety), are unlikely to be fully integrated into models due to the dynamic nature of data and the need for robust security tooling.

Enterprise Adoption: A Hybrid Approach

The future state for enterprise AI will likely involve a blend of frontier closed models and open-source models. The vast majority of tokens (80-90%) will flow through open models, significantly reducing costs and enabling a wider range of use cases. Frontier models will be reserved for the most challenging tasks, with a combination of open and closed models working together for problems in the middle. This mirrors human organizations, where partners delegate work to associates, and cloud computing, where proprietary services are used alongside open-source alternatives.

The emergence of powerful new hardware, such as Nvidia's DGX Spark and Apple Silicon, is making local model execution increasingly viable. These devices can run large language models (20B-40B parameters, sometimes up to 128B) with impressive performance, even rivaling some closed models for specific tasks like coding.

Local vs. Cloud Models: A Complementary Relationship

Olama observes a strong mix of US and Chinese-trained models being used locally, including models from Llama and DeepMind's Gemma. While coding agents often benefit from the power of large cloud models for complex problems, document processing and other straightforward tasks run efficiently and cost-effectively on local hardware. This suggests a hybrid execution model where a router directs tasks to the most appropriate environment, further reducing costs.

Currently, cloud-hosted coding agent models are predominantly Chinese, while local models show a more balanced blend of US, European, and Chinese origins. This highlights a need for more large models from US labs.

The GPU Market and the "Flash" Model Revolution

The GPU market is characterized by rapid price changes and high supply-demand volatility, making access to cutting-edge GPUs challenging for startups. However, inference providers are stepping in to bridge this gap.

A new class of "flash" models, exemplified by DeepSeek Flash, is emerging. These models offer ultra-low cost per token and per task, making them highly efficient for high-volume token usage. They are "good enough" for 80% of tasks, fast, and extremely cheap, enabling widespread adoption and a return to the "unlimited tokens" experience reminiscent of early ChatGPT.

These workhorse models can be chained together to solve complex problems through orchestration, creating new opportunities for startups and existing businesses. This trend suggests a future where orchestration of smaller, specialized models, rather than a single "god model," will be the dominant approach for most customer use cases.

Geopolitics and Model Origin

The geopolitical implications of AI models, particularly their origin, are a growing concern. While some customers prioritize where a model is run and its security, others are deeply concerned about the model's origin due to data provenance and potential biases in its communication style. For mission-critical tasks, understanding the model's origin and ensuring its integrity is paramount.

The "Manchurian candidate" problem, where a model might be subtly booby-trapped, is a concern, though no known cases have been publicly reported. Enterprises, with their robust IT and security teams, are accustomed to managing supply chain risks in open-source software, and with proper safety checks, they believe these issues can be mitigated.

The Olama Journey: From Docker to AI

Olama's co-founders, with backgrounds at Docker, leveraged their experience in building great developer experiences. Their journey involved a period of searching for the right problem to solve, eventually pivoting from Kubernetes security to locally hosted LLMs. This pivot, though initially challenging, proved to be a significant turning point, coinciding with the release of Llama models.

The rapid adoption of Olama, reaching 100,000 GitHub stars much faster than previous projects, demonstrated a clear product-market fit. This success was driven by the fact that open models were free to start with and could run anywhere, appealing to both hobbyists and enterprise IT developers.

Monetization and the Future of Open Source

Olama's monetization strategy evolved, initially focusing on privacy-focused AI products. However, the true opportunity emerged with the rise of coding agents and the ability of open models to service this high-consumption use case. The company recognized the need to wait for the market to mature and for open models to reach a product-level fit comparable to closed models.

The experience highlighted the importance of staying connected with customers and understanding their evolving needs, rather than viewing users as an anonymous "blob on the internet."

The Value of Experience and Community

The founders' decision to join Y Combinator, even as second-time founders, was driven by the desire for community and mentorship. The YC network provided a crucial support system, helping them navigate the challenges of building a startup and learn from the experiences of other founders. This collective wisdom, particularly in avoiding common mistakes and leveraging successful strategies, proved invaluable.

The AI world is breaking many traditional rules of infrastructure and DevOps. Concepts like "platform as a service" being vulnerable are no longer true, as going up the stack can bring startups closer to the customer. Additionally, the inherent non-determinism of LLMs, a "feature not a bug," challenges the traditional systems engineering mindset of perfect design and validation. The ability to build services where no single engineer understands all the code, a reality with AI, also necessitates new approaches to team building and operations.

Olama's role in curating and simplifying the fragmented open model landscape is crucial. By making diverse models, inference technologies, and cloud services work seamlessly, Olama empowers developers to focus on building their applications, rather than navigating a complex and often undocumented ecosystem. This curation, similar to what Open Router and Open Code offer, is becoming increasingly valuable in a world of abundant models and providers.

  Takeaways

  • Enterprises are shifting to open-source AI models to cut costs and gain deeper customization, with platforms like Olama enabling both local and cloud execution for millions of developers.
  • Token consumption has exploded, driven by coding agents such as OpenClaw and Hermes, and Olama reports a 150‑fold increase in open‑model token usage since the start of the year.
  • While fine‑tuning cycles fluctuate, the rapid release of new open models makes custom training harder yet spurs better tooling and a renewed interest in model adaptation.
  • Security concerns remain a barrier, but open‑weight models allow targeted penetration testing and safety training, turning a risk into a strategic advantage for enterprises.
  • The future AI stack will likely be hybrid, with 80‑90% of tokens processed by open models and frontier closed models reserved for the toughest tasks, supported by a growing ecosystem of “flash” low‑cost models and orchestration platforms like Olama.

Frequently Asked Questions

What are 'flash' models and why are they considered sufficient for most AI tasks?

Flash models are ultra‑low‑cost, high‑throughput open‑source LLMs such as DeepSeek Flash that trade a small drop in capability for dramatically cheaper token pricing, making them adequate for about 80 % of typical workloads. They deliver fast inference, keep expenses low, and enable large‑scale token usage without sacrificing core functionality.

How does Olama support rapid 'day zero' model launches for new open models?

Olama supports rapid “day zero” launches by pre‑configuring inference engines, preparing use‑case harnesses, and coordinating with hardware and cloud providers to have a fully optimized stack ready within 24 hours of a model’s release. This playbook ensures models run efficiently, developers get immediate access, and enterprises can quickly evaluate the new model’s performance and cost profile.

Who is Y Combinator on YouTube?

Y Combinator is a YouTube channel that publishes videos on a range of topics. Browse more summaries from this channel below.

Does this page include the full transcript of the video?

Yes, the full transcript for this video is available on this page. Click 'Show transcript' in the sidebar to read it.

Helpful resources related to this video

If you want to practice or explore the concepts discussed in the video, these commonly used tools may help.

Links may be affiliate links. We only include resources that are genuinely relevant to the topic.

Full transcript is not shown on this page

This page focuses on the summary and original notes. For full verification, refer to the original YouTube video.

PDF