Skip to content
Close Menu

    Subscribe to Updates

    Get the latest news from tastytech.

    What's Hot

    Google’s Gemini 3.6 Flash targets enterprise agent token costs

    July 22, 2026

    Agentic AI vs Automation: Key Differences Explained

    July 22, 2026

    VCF 9.1 VPC Networking: Distributed vs. Centralized Transit Gateway Designs

    July 22, 2026
    Facebook X (Twitter) Instagram
    Facebook X (Twitter) Instagram
    tastytech.intastytech.in
    Subscribe
    • AI News & Trends
    • Tech News
    • AI Tools
    • Business & Startups
    • Guides & Tutorials
    • Tech Reviews
    • Automobiles
    • Gaming
    • movies
    tastytech.intastytech.in
    Home»AI Tools»Google’s Gemini 3.6 Flash targets enterprise agent token costs
    Google’s Gemini 3.6 Flash targets enterprise agent token costs
    AI Tools

    Google’s Gemini 3.6 Flash targets enterprise agent token costs

    gvfx00@gmail.comBy gvfx00@gmail.comJuly 22, 2026No Comments5 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email


    Google has released Gemini 3.6 Flash and 3.5 Flash-Lite as new workhorses designed to cut latency and token costs for enterprise AI agents.

    The economics of running autonomous software agents inside a production environment come down to a fixed equation few vendors advertise directly. A model needs to reason through a multi-step task competently, but every extra token it generates while doing so adds cost and delay to a workflow that might run thousands of times an hour.

    Teams building background agents rather than chat interfaces need throughput first and parameter count second. Google’s answer, announced this week, splits that trade-off across three models: Gemini 3.6 Flash for coding and multimodal reasoning, Gemini 3.5 Flash-Lite for high-volume, low-latency work, and a restricted Gemini 3.5 Flash Cyber variant built for vulnerability remediation.

    Table of Contents

    Toggle
      • The math behind Gemini 3.6 Flash
      • Figma, Hebbia, and Harvey put the model to work
      • A cheaper tier for high-volume background agents
      • Gemini 3.5 Flash Cyber: A restricted model for patching code
      • Related posts:
    • Oil tops $116 a barrel as Iran accuses US of preparing invasion | Oil and Gas News
    • AWS's legacy will be in AI success
    • Saudi, UAE, Iraq: Can three pipelines help oil escape Strait of Hormuz? | US-Israel war on Iran News

    The math behind Gemini 3.6 Flash

    Google’s developer documentation for 3.6 Flash centres on one figure: 17 percent fewer output tokens than the prior 3.5 Flash version, based on measurements from the Artificial Analysis Index.

    In specific synthetic tests, including the Datacurve DeepSWE benchmark, Google reports drops in token usage of up to 65 percent. Pricing sits at $1.50/1M input tokens and $7.50/1M output tokens, positioning the model for reasoning loops that run continuously rather than on-demand.

    On DeepSWE, the company records a 49 percent success rate for 3.6 Flash against 37 percent for its predecessor. On MLE Bench, the score moves from 49.7 percent to 63.9 percent, and on Google’s GDPval-AA v2 test – which attempts to measure real-world knowledge work rather than coding puzzles – 3.6 Flash scores 1421 against 1349 for the older model.

    Figma, Hebbia, and Harvey put the model to work

    Figma has integrated 3.6 Flash into its prototyping infrastructure, and according to Matt Colyer, the company’s Director of Product Engineering, the model gives developers a faster route through design iterations without a drop in output quality.

    Legal technology platform Harvey and research tool Hebbia route data through the model for multimodal document work: ingesting raw financial filings, parsing document structure, reading embedded charts, and producing draft reports for review.

    Google also folded a client-side computer-use tool directly into the Gemini API and Gemini Enterprise platforms, removing the custom intermediary software engineers previously built to let models operate on top of an operating system.

    The company reports an OSWorld-Verified score of 83.0 percent, up from 78.4 percent, and says updated safeguards against chemical, biological, radiological, and nuclear misuse improve resistance to jailbreaking without raising refusal rates for benign requests.

    A cheaper tier for high-volume background agents

    Gemini 3.5 Flash-Lite targets a different job: document processing and agentic search running at volume rather than reasoning depth. The Artificial Analysis Index measured the model at 350 output tokens per second, the fastest in the 3.5 series according to Google.

    Pricing runs at $0.3/1M input tokens and $2.5/1M output tokens, cheap enough that engineering teams can route simple, high-volume subagent requests to a minimal thinking level and reserve higher thinking levels for multi-step work.

    On Google’s GDM-MRCR v2 long-context test, Gemini 3.5 Flash-Lite recorded a 72.2 percent success rate against 60.1 percent for its predecessor, and its GDPval-AA v2 score nearly doubled, from 642 to 1140. The model carries the same native computer-use tool as 3.6 Flash.

    Separately, Google says Gemini 3.5 Pro remains in partner testing ahead of a full release, and pre-training for the next Gemini 4 architecture is already underway.

    Gemini 3.5 Flash Cyber: A restricted model for patching code

    Automated vulnerability scanners now surface flaws faster than most security teams can patch them, and that gap is where Google positions Gemini 3.5 Flash Cyber.

    The model is built to validate and remediate code vulnerabilities, and Google reports performance on the CyberGym benchmark competitive with frontier models (though it hasn’t made those figures public in the same detail as its consumer-facing releases.)

    Distribution stays restricted to governments and vetted partners through a pilot programme, a limitation Google frames as a safeguard against the model generating exploit code for offensive use.

    Inside Google’s CodeMender security agent, multiple instances of 3.5 Flash Cyber run in parallel, cross-checking one another’s findings before producing a single remediation report a human reviewer signs off on.

    Engineering teams seeking to integrate these new models can access them through the Gemini API via Google AI Studio, Android Studio, or the Gemini Enterprise Agent Platform. Consumers can also access the new models in the Gemini app and 3.5 Flash-Lite is also rolling out in Google Search.

    See also: Bristol Myers Squibb buys Nvidia AI system for drug discovery

    Banner for the AI & Big Data Expo event series.

    Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

    AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

    Related posts:

    Proving the case on day two at TechEx North America

    White House accuses South Africa of harassing US gov’t staff in latest row | Donald Trump News

    OpenAI Frontier puts enterprise AI agents on a collision course with SaaS

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleAgentic AI vs Automation: Key Differences Explained
    gvfx00@gmail.com
    • Website

    Related Posts

    AI Tools

    Automating VCF 9.1 VPC Networking with PowerCLI: IP Blocks, Subnets, NAT, and External IPs

    July 22, 2026
    AI Tools

    US judge blocks Trump bid to strip work permits from immigrants | Courts News

    July 21, 2026
    AI Tools

    The AI Slot Machine Effect: Why Generative Feeds Disrupt Deep Work And How to Reclaim Focus

    July 21, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Black Swans in Artificial Intelligence — Dan Rose AI

    October 2, 2025212 Views

    Every Clue That Tony Stark Was Always Doctor Doom

    October 20, 2025134 Views

    We let ChatGPT judge impossible superhero debates — here’s how it ruled

    December 31, 2025100 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram

    Subscribe to Updates

    Get the latest tech news from tastytech.

    About Us
    About Us

    TastyTech.in brings you the latest AI, tech news, cybersecurity tips, and gadget insights all in one place. Stay informed, stay secure, and stay ahead with us!

    Most Popular

    Black Swans in Artificial Intelligence — Dan Rose AI

    October 2, 2025212 Views

    Every Clue That Tony Stark Was Always Doctor Doom

    October 20, 2025134 Views

    We let ChatGPT judge impossible superhero debates — here’s how it ruled

    December 31, 2025100 Views

    Subscribe to Updates

    Get the latest news from tastytech.

    Facebook X (Twitter) Instagram Pinterest
    • Homepage
    • About Us
    • Contact Us
    • Privacy Policy
    © 2026 TastyTech. Designed by TastyTech.

    Type above and press Enter to search. Press Esc to cancel.

    Ad Blocker Enabled!
    Ad Blocker Enabled!
    Our website is made possible by displaying online advertisements to our visitors. Please support us by disabling your Ad Blocker.