UPDATED SEPTEMBER 15, 2026
UPDATED SEPTEMBER 15, 2026

The Frontier

Your signal. Your price.

Include
Lookback
||
Nerd Snipe with Theo and Ben
  • · 4d ago

    Theo and Ben argue that Gemini 3.8 Flash is highly inefficient for real-world programming tasks despite its strong performance on benchmarks. Theo points out that the model took over 20 minutes and 100 tool calls to execute a simple four-line code change.

    +3 more
    +1 more
  • · 4d ago

    Ben claims Gemini 3.8 Flash quickly consumed his entire Cursor monthly usage allocation in roughly one hour. Across four development threads, the model processed 1.5 billion tokens and racked up hundreds of dollars in API usage.

    +3 more
    +2 more
  • · 4d ago

    Theo claims he pressured the Gemini and DeepMind teams into adjusting their terms of service regarding anti-gravity support. He initiated this because users risked lifetime Google account bans if their integrations flagged the system.

    +3 more
    +2 more
  • · 4d ago

    Ben suspects that Gemini 3.8 Flash's top ranking on the DeepSWE benchmark is the result of extreme overfitting. He points out that Data Curve, the creator of DeepSWE, makes its money selling training data directly to major labs.

    +3 more
    +3 more
  • · 4d ago

    Ben notes that Muse Spark 1.3 retains the extreme speed and free contributor tier of its predecessor while improving performance. Scale AI's Alexander Wang publicly urged users to test Muse Spark 1.3 Max before judging the model family.

    +3 more
    +2 more
  • · 4d ago

    Theo and Ben argue that OpenAI and Anthropic are cutting profit margins to compete for API dominance. Ben suspects OpenAI's recent price cuts on Soul, Luna, and Terra were side effects of optimizing Astra to match Fable's price.

    +3 more
    +7 more
  • · 4d ago

    Ben notes that cache writes account for approximately 60% of his daily API costs when using Fable 5.1 in Claude Code. He notes that while cache reads are highly cost-effective, cache writing remains an expensive implementation detail.

    +3 more
    +2 more
  • · 4d ago

    Theo theorizes that a one-day delay in the GPT-6 Astra launch was caused by a dispute over cloud infrastructure agreements with AWS. The announcement notably omitted any mention of Azure despite the model's availability there.

    +3 more
    +3 more
  • · 4d ago

    Theo argues that Fable 5.1 consistently produces clean, mergeable code on the first try. In contrast, Astra is prone to over-complicating pull requests by adding unnecessary tests, try-catch blocks, and refusing to delete legacy code.

    +3 more
    +2 more
  • · 4d ago

    Theo utilized Astra to run 40 simultaneous sub-agents overnight to identify and merge performance fixes for T3 Code. The run successfully merged 40 pull requests over seven hours, which would have cost $2,588 at retail API prices.

    +3 more
    +2 more
  • · 4d ago

    Theo claims Fable 5.1 is highly proficient at web design and complex 2D or 3D animations, creating polish unmatched by other models. However, its static layouts can be highly inconsistent and require heavy human guidance.

    +3 more
    +1 more
About The Frontier
End of 7-day results — 11 results
11 results