AI Newsway

Gemini 3.7 Flash Arrives Three Weeks After 3.6 - At Half the Price

Google's new workhorse model posts a 65.3% DeepSWE score and launches at $0.75 per million input tokens

|3 min read0
AI Summary
Google released Gemini 3.7 Flash on Thursday, three weeks after Gemini 3.6 Flash, positioning it as its top workhorse model for coding and agents. It scored 65.3 percent on DeepSWE v1.1, up from 49.0, and 1588 Elo on WebDev Arena, while launching at $0.75 per million input and $3.75 per million output tokens, half the predecessor's price. Rates double to $1.50 and $7.50 on January 1, 2027, giving teams four and a half months to build.
A silhouette rendered from streaming code, reflecting Gemini 3.7 Flash's focus on software engineering and agent workloads
A silhouette rendered from streaming code, reflecting Gemini 3.7 Flash's focus on software engineering and agent workloads

Google has released Gemini 3.7 Flash, positioning the model as its most capable workhorse tier for coding and agent workloads. The launch arrived on Thursday, only three weeks after Gemini 3.6 Flash reached developers, and Google is pairing it with an introductory price that undercuts its predecessor by half.

The compressed release cycle is unusual even by the standards of the current model race. Google attributes the turnaround to a combination of direct developer feedback and algorithmic work that the company says will carry forward into future releases. Notably, the model that shipped is not the Gemini 3.5 Pro that many developers had been anticipating.

Benchmark Gains Concentrated in Software Engineering

The clearest improvements show up in code. On DeepSWE v1.1, an evaluation of issue resolution and debugging, Gemini 3.7 Flash scored 65.3%, up from 49.0% for the 3.6 release. FrontierCode 1.1 Main climbed to 43.6% from 34.4%, a benchmark Google uses to measure first-pass accuracy on production-ready code.

Web development results moved in the same direction. The model posted an Elo score of 1588 on Arena.ai's WebDev Arena, ahead of 3.6 Flash at 1538. Google says the model produces more feature-complete applications in fewer prompts and holds closer to a reference design, whether that reference is a screenshot or a full design system.

Knowledge-heavy domains saw the largest proportional jumps. On GDP.pdf, which tests document processing in fields such as finance, law, and biosciences, the score rose to 34.0% from 22.0%. AutomationBench, which measures completion of real business workflows, nearly doubled to 30.4% from 17.0%.

Pricing Is the Headline for Developers

Through December 31, 2026, Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens. On January 1, 2027, those rates double to $1.50 and $7.50 respectively. The introductory window effectively gives teams four and a half months to build against the cheaper tier before standard pricing applies.

That combination of a measurable intelligence gain alongside a 50% price cut, delivered in roughly three weeks, is what drew the most developer attention following the announcement. For teams running agents at production volume, per-token cost tends to dominate the deployment calculus more than headline benchmark numbers.

Availability and Agent Behavior

Google describes behavioral changes that matter for autonomous workloads. The company says the model adapts more readily when it hits a roadblock, asks for clarification when intent is ambiguous, and invests more effort in multi-step planning and tool calls. The practical claim is fewer retries and less manual oversight.

Developers can reach the model through the Gemini API via Google AI Studio, Android Studio, and the agent-first Google Antigravity environment. Enterprise customers get access through the Gemini Enterprise Agent Platform and the Gemini Enterprise app.

Consumers encounter it through Gemini Spark, the 24/7 personal agent Google introduced at I/O, which switches to 3.7 Flash starting today for Google AI Pro and Ultra subscribers across more than 160 countries. Google says the update sharpens Spark's tool use across Workspace apps for multi-step tasks like consolidating files and drafting emails.

The release ships with updated Frontier Safety safeguards covering chemical, biological, radiological, and nuclear misuse as well as cyber offense. With the next Pro-tier model still outstanding, the pace of Google's Flash cadence now sets the tempo competitors have to answer.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Qwen3.8 Max Tops Artificial Analysis Agentic Index, Outranking Every US Lab but Two
LLM & Chatbots

Qwen3.8 Max Tops Artificial Analysis Agentic Index, Outranking Every US Lab but Two

Alibaba's Qwen3.8 Max leads the Artificial Analysis agentic index, winning through long-horizon persistence rather than top reasoning scores.

Seung Jung41 days ago
GLM-5.3 Scores 60 on Artificial Analysis Index at Half the Usual Output Price
LLM & Chatbots

GLM-5.3 Scores 60 on Artificial Analysis Index at Half the Usual Output Price

Artificial Analysis scored Z.ai's GLM-5.3 at 60 on its Intelligence Index, well above the 35 median, at $4.40 per million output tokens. The catch is verbosity.

Seung Jung29 days ago
DeepSeek V4 Pro Hits General Availability at a Fraction of Frontier Pricing
LLM & Chatbots

DeepSeek V4 Pro Hits General Availability at a Fraction of Frontier Pricing

DeepSeek has moved its V4 Pro model to general availability, and the release is drawing attention less for raw capability than for what that capability now cost...

Seung Jung34 days ago
GPT-6 Astra Goes to Work: OpenAI's Priciest Model Bets Everything on Computer Use
LLM & Chatbots

GPT-6 Astra Goes to Work: OpenAI's Priciest Model Bets Everything on Computer Use

OpenAI has begun rolling out GPT-6 Astra to business customers, framing its newest frontier model less as a chatbot and more as a worker that operates software...

Seung Jung7 days ago
Grok 4.6 Reaches the AI Frontier Without Raising Its Price
LLM & Chatbots

Grok 4.6 Reaches the AI Frontier Without Raising Its Price

xAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol while holding pricing flat at $2/$6 per million tokens.

Seung Jung34 days ago
Grok Bot Lets AI Agents Run Their Own Group Chat — and Sign Into Your Accounts
LLM & Chatbots

Grok Bot Lets AI Agents Run Their Own Group Chat — and Sign Into Your Accounts

SpaceXAI opened a Grok Bot beta where multiple agents coordinate in group chats, assign ownership to each other, and sign into a user's own accounts.

Seung Jung33 days ago