Matthew Berman

GOOGLE IS BACK! (Gemini 3.8 Flash)

Sep 3, 2026 17 min
gemini 3.8 flashai benchmarkslarge language models
Watch on YouTube Follow Matthew Berman on Rundown — free

Summary

AI summaries can be incomplete or wrong. Verify anything important against the original video.

This review analyzes the capabilities and performance of Google's new Gemini 3.8 Flash AI model compared to competing models from OpenAI and Anthropic.

The video reviews Google's recently released Gemini 3.8 Flash model, evaluating its performance across various industry-standard benchmarks including software engineering tasks (DeepSWE), data analysis (GDP), legal reasoning, and agentic coding capabilities. The host highlights that while the model is not the absolute top-tier performer in every metric, it delivers high-quality results at a significantly lower price point, making it highly competitive and cost-effective for enterprise use. In addition to quantitative benchmarks, the host presents subjective testing involving the generation of 3D dioramas, web page design, and a PowerPoint-style presentation, noting mixed results in creative and design-oriented tasks. The review concludes by emphasizing that the model's value proposition lies in its balance of capability and extreme affordability, especially when compared to more expensive alternatives like Claude Opus or GPT-5.6 models.

Verdict

Gemini 3.8 Flash
ai model · $0.75 per million tokens input, $3.75 per million tokens output

Gemini 3.8 Flash is an exceptionally high-value model that offers impressive performance at a fraction of the cost of top-tier competitors.

Depends

Pros

  • Very low cost per token compared to major competitors. 5:40
  • Strong performance on software engineering (DeepSWE) and legal benchmarks. 3:31
  • Excellent performance on specialized cyber-security vulnerability tasks. 13:55

Cons

  • Lower accuracy on some real-world knowledge benchmarks (GDP) compared to Claude Opus 5. 4:09
  • Inconsistent performance in creative tasks like web design or 3D diorama generation. 12:28

Specs

DeepSWE v1.1 benchmark 73.7% 3:30
CyberGym Pass@1 86.2% 14:17

Compared to

  • Claude Opus 5

    More expensive but generally outperforms in high-level reasoning and complex knowledge work.

  • GPT-5.6 Sol

    Offers comparable or slightly inferior results to Gemini 3.8 Flash in many tasks at a significantly higher price.

Best for

  • Cost-conscious enterprise applications
  • High-volume software engineering tasks
  • Security-focused vulnerability scanning

Not for

  • Users requiring the absolute highest performance for complex creative tasks
  • Users who only prioritize design consistency

Key Points

  • 2:13 Detailed breakdown of benchmark scores: Gemini 3.8 Flash, Claude Opus 5, and GPT-5.6 models.
  • 2:14 Analysis of DeepSWE and GDP benchmark performance, focusing on long-horizon software and knowledge work.
  • 5:32 Pricing comparison analysis illustrating the cost-advantage of Gemini 3.8 Flash.
  • 7:29 Evaluation of legal reasoning (Harvey's Legal Benchmark) and agentic terminal coding performance.
  • 13:27 Overview of Gemini 3.8 Flash Cyber, a specialized model for vulnerability discovery.
  • 16:35 Subjective demos including 3D diorama generation, web design, and presentation creation.
  • Introduction to Gemini 3.8 Flash and the current competitive landscape of AI models.

Worth watching if: You are a software developer, enterprise decision-maker, or AI enthusiast trying to decide whether to integrate Gemini 3.8 Flash into your workflows based on cost and performance trade-offs.

Get every Matthew Berman video extracted like this

One daily email with structured extracts of every channel you follow. Free tier covers 15 videos a month.

Sign in with Google

No credit card. Free tier forever.

Watch on YouTube