Google has updated Gemma 4, their open-weight multi-modal LLM with 256K context length supporting text, video, audio, and code, featuring five key improvements: reliable structured tool calling with 8-point improvement on airline tasks and 6-point improvement on telecom tasks, optimized vision token budget allowing developers to trade inference cost against image detail, flash attention 4 with Hopper GPU speedup for faster inference and lower memory usage, cleaner chat templates reducing formatting errors and model laziness, and more complete answers with improved reasoning capabilities for production-grade AI agents, customer support bots, coding assistants, and document processing applications.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Google New Gemma 4
Added:So, hi everyone. Welcome back to Data Science in a pocket. And while the Gemini 1.5 storm is going on the internet, Google has very silently updated Gemini 4, the best small LLM, with some major updates in tool calling, vision budget, flash tension 4, and less laziness. It has the same weights, but the major updates will obviously impact your production-grade results. So, let's get started. Let's try to understand how to use the new Gemini 4 and what are its key features.
What is Gemini 4? First of all, let's try to understand that. Google's open weight family of multi-LLMs, multi-model LLMs built for text, images, reasoning, and agent workflows, which is open source. It has a context length of 256K, has four modalities, text, video, audio, code, and open weights and license.
There are five improvements that our team has done in the new updated Gemini 4. One is tool calling. Reliable structured function calling. It has improved on its agentic capabilities, I would say. Vision token budget. Tune image details versus cost. They have optimized it for flash attention [snorts] 4. Hopper GPU speedup. Chat template cleanup prompt formatting and reduced laziness. More complete answers.
Structured calls finally reliable. So, let's first of all understand what does it has improved on each aspect one by one. Previously, Gemini 4 sometimes returned plain text instead of structured tool calls in long conversation. This has been improved.
Correct structured function calls consumes return tool output clearly. And two-step agent workflow is verified.
Choose how much detail to extract.
Developers now control the vision token budget. Trade inference cost against image details.
Fast and cheap. If you are working with low budget, aggressive image resize, low token count, fast inference, lower cost.
But if you want high quality results, you need to tune into high budget for vision. Maximum detail preserves fine details, better OCR on small text, read charts, tables, and diagrams.
Hopper class speed up. Optimized attention algorithm tuned specifically for NVIDIA Hopper GPU. Faster inference, lower memory, better throughput. So, when they have improved implemented flash attention for, it has become faster inference latency, higher GPU throughput. It will generate more tokens on the same GPU, lower memory usage, and better utilization. As you can see what are the GPUs that would get affected, H100, H200, and GH200.
Older NVIDIA architecture won't get impacted by this.
Cleaner templates, less lazy answers.
That's a major improvement. Improved chart template. Updated official template reduces formatting errors and improves capabilities with VLM and other engines and reduce model laziness. Gemma 4 now produces more complete answers.
Earlier, it was trying to just complete everything, but now it is focusing more on detailing as well, especially for reasoning-heavy workloads.
Agentic tool call measured. Google's published benchmark delta for tool calling task. It has increased by eight points on airline tool calling and by six points on telecom tool calling. Uh these are Google's published benchmarks.
Tested with VLM ready for production.
Deployment use the Gemma 4 tool parcel, reasoning parcel, and updated chart template also. Who will benefit from the updates? AI agents, customer support bot, coding assistants, document processing, OCR, and API-connected assistants as well.
Yes, update today. Same weights, better everything.
The model has an open source also, and I will just show you how to use the model on Hugging Face well.
When you browse to Hugging Face, you don't have to do much. Just search for Gemma 4, and it's the same old models the team has updated. As you can see, updated 4 days back. So, all the Gemma 4 existing models has been updated with the new update. So, you don't need to do anything else. Whatever you're working with, it has automatically improved. If you have already installed it, I would suggest to update that model, too. With this, it's a wrap. I hope you try out the new Gemma 4, and let me know in the comment section how you feel about the model. It's the best small LM I would again suggest. Thank you so much.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

SuperBike Factory Has Gone... What's Next for the Motorcycle Industry?
thatbikersimon
11K views•2026-07-22