HOME>Blog>Google launches Gemini 3.7 Flash for coding and agents
AI News5 min read

Google launches Gemini 3.7 Flash for coding and agents

Google has introduced Gemini 3.7 Flash, a workhorse model focused on coding, knowledge work, web development, and agent workflows, with introductory API pricing through the end of 2026.

Kylon TeamProduct

The short version

Google has introduced Gemini 3.7 Flash, describing it as its most intelligent workhorse model yet for coding and agents. The release arrives three weeks after Gemini 3.6 Flash and targets software engineering, knowledge work, web development, and business workflow automation. Google's official announcement is the source for the launch date, positioning, and benchmark figures below.

The commercial headline is the introductory price: $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026. Google says this is half the original Gemini 3.6 Flash cost per million tokens.

What changed from Gemini 3.6 Flash

Google says Gemini 3.7 Flash improves coding tasks such as debugging and issue resolution, first-pass code accuracy, production-ready code generation, web development, document comprehension, and enterprise workflow automation. These are Google's reported evaluations, not independent verification. See the full evaluation details.

The official announcement reports the following comparisons with Gemini 3.6 Flash:

  • FrontierCode 1.1 Main: 43.6% versus 34.4%.
  • DeepSWE v1.1: 65.3% versus 49.0%.
  • WebDev Arena: Elo 1588 versus 1538.
  • GDP.pdf: 34.0% versus 22.0%.
  • AutomationBench: 30.4% versus 17.0%.

Google's launch post contains the benchmark names, scores, and evaluation links. Teams should reproduce the tests with their own data before using the numbers for a model decision.

Pricing and availability

Gemini 3.7 Flash is available through the end of 2026 at $0.75 per million input tokens and $3.75 per million output tokens, according to Google. Developers can access it through Google Antigravity, Gemini API in Google AI Studio, and Android Studio. Enterprises can use it through Gemini Enterprise Agent Platform and Gemini Enterprise. Availability and pricing.

Google also says Gemini Spark will use Gemini 3.7 Flash starting on launch day for Google AI Pro and Ultra subscribers in supported countries. Google's announcement describes Spark as a personal agent that can work with Google Workspace apps.

The announcement does not give a general replacement date for Gemini 3.6 Flash. Teams should check the current API documentation and model availability in their own region before planning a migration.

How it compares with OpenAI's latest speed push

Google and OpenAI are emphasizing different levers. Gemini 3.7 Flash focuses on a workhorse model with improved coding and agent workflows, paired with introductory token pricing. OpenAI's GPT-5.6 Sol Ultrafast preview keeps GPT-5.6 Sol as the underlying model and adds a processing tier that OpenAI says can run up to 14 times faster than Standard processing, with up to 750 output tokens per second. OpenAI's official announcement provides those speed and preview details.

For a team building an agent, the choice is not simply Google versus OpenAI. A high-volume extraction or routine workflow may benefit from Gemini 3.7 Flash's introductory token economics. A live interaction may place more value on latency. In both cases, the meaningful unit is a completed task with its tool calls, retries, review steps, and downstream effects.

What developers should test

A practical evaluation can use a small, representative set of tasks:

  1. Debugging a real issue with the same repository context used in production.
  2. Generating a web interface from an existing design system or screenshot.
  3. Extracting decisions and obligations from long business documents.
  4. Completing a multi-step workflow with tools, structured output, and a human approval step.
  5. Measuring retries, correction time, and total cost per completed task.

Google says Gemini 3.7 Flash better adapts to roadblocks, clarifies intent when needed, follows instructions more closely, and makes more deliberate tool calls. These claims come from Google's announcement, so they should be validated against the team's own workflow.

Why the workflow around the model matters

A model upgrade is easier to evaluate when the request, source documents, tool permissions, human review, and final record are connected. Otherwise, teams may compare isolated chat responses while missing the operational work around them.

Kylon gives teams a shared workspace for channels, files, agents, connections, skills, and recurring workflows. That makes it possible to route different steps to different models while preserving the same business context and review path. The model is one part of the system. The surrounding work structure determines whether the result can be used repeatedly.

Bottom line

  • Gemini 3.7 Flash is Google's new workhorse model for coding and agents. Official source
  • Google reports improvements across coding, web development, document comprehension, and enterprise workflow automation. Official source
  • Introductory API pricing is $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026. Official source
  • Compare the model with real tasks, tool calls, retries, review time, and total completed-task cost.

For the full model details and access paths, read Google's official Gemini 3.7 Flash announcement.

Hire the AI agent team that runs your entire business.