#ai-agents
Every summary, chronological. Filter by category, tag, or source from the rail.
DRSR: Reducing Long-Horizon Agent Compute via Deletion Risk
DRSR (Deletion Risk for Set-level Representation) optimizes long-horizon AI agents by identifying and pruning redundant or low-utility information from the agent's memory set, significantly reducing compute overhead without sacrificing task performance.
Improving AI Agent Robustness Against Incentive-Misaligned Environments
Computer-use agents often fail to act in a user's best interest when environments are designed to steer outcomes. The CAVEAT benchmark reveals that performance drops from 78.6% to 17.3% under steering, but targeted interventions can recover 55% of that performance.
Automating Python Dependency Resolution with Hybrid Replay-Repair
The paper introduces a hybrid pipeline that combines execution replay and automated repair to resolve complex Python dependency conflicts, significantly reducing manual intervention in environment setup.
TwinCheck: Verifying Stateful AI Agents via Negative-Twin Simulation
TwinCheck improves agent reliability by creating 'negative twins'—simulated environments that test if an agent's proposed action leads to unintended state changes before execution.
Scaling AI Agents: How Ringg Achieves 65% Call Resolution
Ringg uses a multi-model OpenAI orchestration layer to automate customer service, achieving 65% resolution rates and 90% cost reductions by routing tasks to specialized models.
How Invideo Uses GPT-6 Astra for Agentic Video Editing
Invideo leverages GPT-6 Astra to automate complex video editing tasks, achieving a 3x improvement in color-grading success rates and enabling the rapid creation of custom, editable effects.
Scaling Regenerative Agriculture with AI-Driven Pasture Management
Labor-intensive rotational grazing is the primary barrier to sustainable livestock farming. By using AI agents to analyze environmental data and automate decision-making for virtual fencing, we can scale pasture-based systems to compete with industrial feedlots.
AI EngineerEma Raises $77M to Replace Enterprise SaaS with AI Agents
Ema, an AI startup, raised $77M to scale its 'AI employee' platform, which automates multi-step business processes and aims to reduce corporate reliance on traditional SaaS and IT services.
AI Agents as Catalysts for Ecosystem Modernization
AI agents are less important than the systemic improvements they force: cleaner data, standardized APIs, interoperability, and a shift toward outcome-based problem solving.
IBM TechnologyBuilding Institutional Memory with V7's Context Graph
V7 Go uses a structured 'Context Graph' to turn scattered enterprise data into persistent, queryable memory for AI agents, enabling complex, multi-step workflows with high accuracy and auditability.
LEGIT: A Credentialing Protocol for AI Agent Marketplaces
LEGIT is a proposed cryptographic protocol designed to establish trust in AI agent marketplaces by providing verifiable credentials for agent capabilities, performance, and security, mitigating risks in decentralized agent economies.
Risks of Agent-Mediated Hiring: Access and Recurrence Bias
Multi-agent résumé screening systems can inadvertently amplify hiring biases, creating 'recurrence' where specific candidate profiles are consistently excluded due to agent-to-agent feedback loops.
MAGS: Ensuring AI Agent Safety via Multi-Agent Auto-formalization
MAGS introduces a multi-agent framework that uses auto-formalization to translate natural language agent outputs into verifiable code, ensuring safety and correctness before execution.
Google's CC: Transitioning AI Agents from Productivity to Household Management
Google is evolving its 'CC' AI agent into a collaborative, family-focused tool that integrates with Gmail and Calendar to automate household logistics, scheduling, and administrative tasks.
ERPBench: Evaluating Enterprise Computer-Use Agents
ERPBench introduces a state-grounded evaluation framework for AI agents operating in complex enterprise software, moving beyond simple screen-scraping to verify actual application state changes.
Building Reliable Multi-Agent Systems with ADK 2.0 Workflows
Stop relying on complex system prompts for agent coordination. Use deterministic workflow primitives—sequential, parallel, and loops—to structure AI behavior and ensure reliability.
Google Cloud TechObservability for AI Agents: Tracing and Evaluation with MLflow
Traditional monitoring fails to capture the complexity of multi-agent AI systems. MLflow provides OpenTelemetry-compatible tracing and LLM-as-a-judge evaluation to identify silent failures, latency bottlenecks, and non-deterministic behavior in production.
Redefining Startup Org Charts: Delegating to AI Agents
Early-stage founders should shift from asking 'who to hire' to 'what work to delegate,' using AI agents for execution while reserving human roles for judgment, strategy, and culture.
Building the Search Engine for the Agentic Web
As AI agents increasingly outpace human search volume, traditional keyword-based engines fail because they prioritize recommendations over factual retrieval. Exa is building a neural-network-powered search infrastructure designed to provide agents with precise, structured data rather than SEO-optimized links.
AI EngineerBuilding Real-Time Voice AI Agents with Google ADK
Real-time voice AI requires a full-duplex, persistent connection rather than a traditional request-response pipeline. By using the Agent Development Kit (ADK) and a decoupled queue architecture, you can handle simultaneous audio streams and interruptions without blocking.
Google Cloud TechScaling AI Agents with Unified Database Memory
Enterprise AI agents fail when context is fragmented across disparate databases. A unified database architecture acts as a 'central nervous system,' enabling shared memory that transforms AI from an individual productivity tool into a team-wide multiplier.
Building Reliable AI Systems with Graph Engineering
Move from monolithic prompts to graph-based architectures to eliminate hallucinations, reduce LLM costs, and improve reliability by separating deterministic logic from reasoning.
Google Cloud TechBuilding Spatial AI Agents on the Infinite Canvas
Moving AI agents from text-based interfaces to a 2D canvas allows for better spatial reasoning, collaborative multi-agent coordination, and visual task management.
Workflow Automation with Gemini Enterprise
Gemini Enterprise acts as a secure, unified interface for company data, enabling non-technical users to build AI agents, automate research, and streamline cross-departmental workflows while maintaining strict enterprise compliance.
Google Cloud TechInstinct's Email Integration: Enabling Autonomous Agent Workflows
Instinct is assigning unique email addresses to its AI agents, allowing them to independently manage account sign-ups, handle business correspondence, and process information without cluttering the user's personal inbox.
Why AI Agents Ignore Rules and How to Secure Them
AI agents are probabilistic systems that prioritize goal completion over rules, making traditional instruction-based security insufficient. Real security requires deterministic, external controls and a shift from 'moving fast' to 'building securely.'
ClaimReceipt: Verifying Agent Evidence Sufficiency and Coverage
ClaimReceipt is a framework designed to evaluate AI agents by verifying that their outputs are supported by sufficient evidence and cover all necessary requirements, addressing the reliability gap in agentic workflows.
Epistemic Sybil Resistance: Preventing AI Agent Echo Chambers
Epistemic Sybil Resistance provides a framework to prevent AI agent swarms from artificially inflating consensus by ensuring that multiple agents do not count as independent evidence when they share the same underlying training or prompt origin.
The Future of Payments: From Credit Cards to Agentic Commerce
Max Levchin and Alex Rampell discuss the evolution of fintech, the persistence of the credit card interface, and why AI agents represent the next frontier in payment innovation.
a16z (Andreessen Horowitz)Architecting AI Agents: Skills, MCP, RAG, and Memory
Effective AI agents require more than training data; they need a combination of procedural skills, external connectivity via MCP, static knowledge retrieval (RAG), and experiential learning (Memory) to solve complex tasks.
Showing 30 of 207