Which AI Models Are Worth Using
Summary
AI summaries can be incomplete or wrong. Verify anything important against the original video.
This video presents a personal tier list of current AI models, evaluating them based on performance, cost-efficiency, and practical utility for coding and agent-based workflows.
The video provides a detailed breakdown of a user's experience with various LLMs, categorized into a tier list structure. The creator focuses on real-world utility for software development, specifically looking at how these models perform in coding tasks and agentic environments, including their ability to access the internet and their token efficiency. Throughout the session, the creator tests different models, including Fable 5, GPT-5-6-sol, Kimi K3, and various Gemini models, while justifying their rankings based on performance, cost per task, and specific features like vision capabilities.
The review emphasizes that the value of an AI model is not strictly tied to speed or benchmark scores. Instead, the creator prioritizes reliability in long-running tasks, cost-effectiveness for high-volume use cases, and the ability to integrate into professional coding environments like Cursor or T3 Code. The video also touches on the importance of licensing and the impact of token usage on real-world costs, providing a practical look at how different AI providers, such as OpenRouter, structure their offerings.
Verdict
There is no single 'best' model; choice should be driven by specific task requirements such as cost-efficiency, speed, and specialized capabilities like vision or agentic planning.
Pros
Cons
Specs
| Pricing structure | Varies by model and provider (e.g., $1.20 - $15.00 per million output tokens). | 8:16 |
Best for
Not for
Key Points
- 1:42 Establishing a tier list with GPT-5-6-sol as the baseline.
- 2:11 The importance of internet access for AI models in coding tasks.
- 7:28 Review and tier placement of GPT-5-6-terra.
- 11:40 Review and tier placement of GPT-5-6-luna (tier A).
- 13:30 Review of DeepSeek v4 models (Flash and Pro).
- 19:57 Review of Kimi K3, highlighting performance and pricing.
- 28:35 Review of Sonnet 5, placed in D tier.
- 31:00 Finalizing the S tier with Fable 5.
- 35:06 Review of Gemini Flash and Pro models, questioning Google's direction.
- Introduction and criteria for AI model evaluation.
Worth watching if: You are a software developer trying to decide which AI model to use for your coding tasks, or you are interested in a practical, hands-on comparison of LLMs beyond generic benchmarks.
Get every Theo - t3․gg video extracted like this
One daily email with structured extracts of every channel you follow. Free tier covers 15 videos a month.
Sign in with GoogleNo credit card. Free tier forever.