Google Cloud Tech
Every summary, chronological. Filter by category, tag, or source from the rail.
Building Production-Ready Apps with Gemini 3.5 Transcribe
Gemini 3.5 Transcribe offers two distinct APIs for speech-to-text: synchronous batch processing for pre-recorded files and the Live API for real-time streaming, both supporting advanced features like diarization, word-level timestamps, and custom vocabulary.
Google Cloud TechBuilding AI-Powered Transcription Pipelines with Gemini 3.5
Gemini 3.5 Transcribe enables developers to build high-accuracy, domain-specific transcription pipelines for both live and batch audio without requiring model training.
Preventing Surprise Cloud Bills with Hard Spending Caps
Google Cloud allows developers to set hard spending caps on specific projects for services like Gemini API and Vertex AI, automatically disabling resources when a budget threshold is reached to prevent runaway costs.
Implementing Real-Time Tool Calling for Voice AI Agents
To build responsive voice agents, use a synchronous tool-calling loop where the model decides and your code executes. Keep tools instant to avoid conversation gaps and use a 'before_tool_callback' checkpoint to enforce policies and handle slow actions.
Google Cloud TechBuilding Reliable Multi-Agent Systems with ADK 2.0 Workflows
Stop relying on complex system prompts for agent coordination. Use deterministic workflow primitives—sequential, parallel, and loops—to structure AI behavior and ensure reliability.
Google Cloud TechBuilding Real-Time Voice AI Agents with Google ADK
Real-time voice AI requires a full-duplex, persistent connection rather than a traditional request-response pipeline. By using the Agent Development Kit (ADK) and a decoupled queue architecture, you can handle simultaneous audio streams and interruptions without blocking.
Google Cloud TechMastering Agent Harnesses: The Stack Behind Autonomous Coding
An agent harness is the essential infrastructure wrapping an LLM that provides the tools, memory, and guardrails necessary to transform raw model capabilities into reliable, autonomous software engineering workflows.
Implementing BM25 for Hybrid Search in AlloyDB & Cloud SQL
Google Cloud has added native BM25 support to AlloyDB and Cloud SQL, enabling high-quality, industry-standard full-text ranking directly within the database to improve RAG and hybrid search performance.
Building Real-Time AI Agents for High-Stakes Environments
The Google Antigravity team built a real-time AI race coach by solving for extreme latency and offline constraints, demonstrating how agent-first development platforms enable rapid iteration in high-stakes, physical environments.
Google Cloud TechBuilding Reliable AI Systems with Graph Engineering
Move from monolithic prompts to graph-based architectures to eliminate hallucinations, reduce LLM costs, and improve reliability by separating deterministic logic from reasoning.
Workflow Automation with Gemini Enterprise
Gemini Enterprise acts as a secure, unified interface for company data, enabling non-technical users to build AI agents, automate research, and streamline cross-departmental workflows while maintaining strict enterprise compliance.
Google Cloud TechGraph Engineering for Predictable AI Workflows
Graph engineering provides a structured, deterministic approach to building multi-agent systems by defining explicit nodes and edges, offering superior control and debuggability compared to agent swarms or simple loops.
Google Cloud TechObserving and Scaling Gemini Agents with Grafana Cloud
Scale AI agents from local testing to production by using the Grafana Sigil SDK to capture telemetry, enabling automated performance monitoring, tool-call inspection, and AI-driven remediation workflows.
End-to-End Agentic Development with Gemini and GitLab
By integrating Google Gemini with the GitLab Duo Agent Platform and Antigravity IDE, developers can automate the entire lifecycle of a feature—from UI design and issue tracking to code generation, automated security reviews, and cloud deployment.
Google Cloud TechBuilding Production-Ready RAG Agents on Google Cloud
Learn to build and deploy a secure, grounded RAG agent using the Google Agent Development Kit (ADK), Streamlit, and Cloud Run, moving from local prototypes to enterprise-ready infrastructure.
Google Cloud TechMoving AI Beyond Code Generation: The Production-Grade SDLC
Coding agents have solved the initial creation of code, but the real bottleneck is production operations. True AI-driven engineering requires organizational memory, live system context, and proactive monitoring to prevent regressions before they trigger alerts.
Google Cloud TechRapid Prototyping and Deployment with Google AI Studio
Use Google AI Studio's build mode to generate, iterate, and deploy full-stack web applications via natural language prompts, bypassing manual coding for initial scaffolding.
Google Cloud TechBuilding and Deploying Full-Stack AI Apps with Firebase
Learn to build, secure, and deploy a real-time, full-stack to-do application using Google AI Studio and Firebase, leveraging automated authentication and real-time database synchronization.
Building and Deploying Turn-Based Web Games with AI
Learn to build real-time, turn-based web games using event sourcing, Firestore for state synchronization, and Google AI Studio for iterative debugging and deployment.
7 Modular Design Patterns for AI Coding Agents
Improve AI coding agent performance by replacing long, confusing prompts with modular 'skills'—specialized text files that the agent loads dynamically only when needed.
Google Cloud TechBuilding Real-Time Voice AI Agents with Gemini Live
Gemini Live enables bidirectional, audio-native conversations by using WebSockets for streaming and built-in voice activity detection to handle interruptions and tool execution.
Google Cloud TechBuilding and Scaling Multi-Agent AI Systems on GKE
A practical guide to deploying AI agents on GKE, using the Model Context Protocol for infrastructure troubleshooting, and implementing secure sandboxing for AI-generated code.
Strategies for Serving JAX Models in Production
Moving JAX models from notebooks to production requires choosing the right serialization and compilation strategy to avoid latency spikes caused by just-in-time compilation.
Scaling JAX Models to Multi-GPU Systems
Scale JAX models across multiple GPUs by defining array layouts with Mesh and PartitionSpec, allowing the compiler to handle gradient synchronization automatically.
Building and Optimizing JAX Training Loops
Build high-performance JAX training loops by maintaining pure functions, keeping data on-device, and utilizing fused kernels like cuDNN attention to avoid GPU memory bottlenecks.
Optimizing JAX Performance on NVIDIA GPUs
JAX performance hinges on ensuring your code runs on the GPU, maintaining stable input shapes to prevent re-compilation, and correctly handling asynchronous execution during profiling.
4 Common Loop Engineering Failures and How to Fix Them
Loop engineering automates repetitive tasks by setting goals and retrying, but it often fails due to runaway costs, confirmation bias, vague objectives, or excessive complexity. Success requires strict stop rules, external evaluation, concrete metrics, and transitioning to graph-based architectures for complex workflows.
Google Cloud TechModernizing Legacy Codebases with AI Agents
Tackle legacy code by treating AI as a coworker: use a three-step 'plan, execute, verify' workflow, prioritize test-driven development, and enforce strict guardrails to prevent hallucinations and errors.
Google Cloud TechBuilding Bidirectional Multimodal AI Agents
Moving from turn-based chatbots to 'omni-apps' requires a continuous loop of perception, reasoning, and expression that handles real-time voice, vision, and browser interaction.
Google Cloud TechBuilding AI Agents with Gemini Enterprise & Google Workspace
Learn how to integrate Gemini Enterprise agents with Google Workspace data and actions using connectors, MCPs, and no-code/pro-code development frameworks to automate enterprise workflows.
Google Cloud TechShowing 30 of 157