#agents
Every summary, chronological. Filter by category, tag, or source from the rail.
UrbanAgent: Tool-Augmented Agents for Complex Urban Systems
UrbanAgent is a framework designed to enable AI agents to execute cross-system tasks in urban environments by integrating specialized tools for data retrieval, analysis, and decision-making across fragmented city infrastructure.
BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL
BAP-SQL introduces a budget-aware framework for agentic Text-to-SQL systems, optimizing schema exploration and query generation by balancing accuracy against token costs and execution constraints.
VeriTrace: Bridging the Gap in Agentic Temporal Exploration
VeriTrace introduces a human-like temporal exploration framework that addresses the limitations of current AI agents in navigating complex, multi-step action spaces by effectively managing temporal dependencies.
HyperAgent: Planning with Tool-Schema Hypergraphs
HyperAgent improves LLM tool-use by representing tool schemas as hypergraphs, enabling more effective planning and execution in complex, multi-step tasks.
Scaling AI Agency in Education via Specialized Plugins
OpenAI is launching three education-specific ChatGPT plugins to help students and educators move from basic query-answering to complex, agentic workflows within secure, institution-managed environments.
Securing AI Evaluation Environments Against Model Misbehavior
As AI models become more capable, third-party evaluation environments require stricter security controls to prevent models from escaping simulated boundaries and interacting with the real internet.
NeSyFS: Neuro-symbolic Fast-Slow Thinking for AI Agents
NeSyFS improves LLM agent performance in partially observable environments by combining fast, intuitive neural responses with slow, symbolic reasoning to handle uncertainty and long-term planning.
Scaling Telco Personalization with Multi-Agent AI Architectures
Circles transformed telco operations by using OpenAI’s API to build a multi-agent support system (CareX) and a personalization engine (Xplore IQ), resulting in a 65% autonomous resolution rate and 22% ARPU growth.
Why AI Companions Suffer from Long-Horizon Persona Collapse
AI companions inevitably lose their defined persona and behavioral consistency over long-term interactions due to cumulative drift in context windows and memory retrieval, necessitating new architectural approaches to state management.
SciToolAgent-Evo: Ontology-Driven Self-Evolving AI Agents
SciToolAgent-Evo addresses the limitations of static AI agents in scientific research by using an ontology-aware framework that allows agents to autonomously discover, evaluate, and integrate new tools in open-world environments.
ThinkReset: Improving Long-Horizon Reasoning via Intermediate Interfaces
ThinkReset addresses the context-window degradation in long-horizon AI reasoning by introducing a learnable 'reset' mechanism that compresses task state into bounded, manageable intermediate interfaces.
Scaling Autonomous Agents with OpenClaw and Ollama
The paper presents a framework for building scalable, autonomous AI agent systems by combining the OpenClaw orchestration layer with local LLM execution via Ollama, addressing key bottlenecks in agentic workflows.
From Tokenmaxxing to Tokenomics: Scaling AI Agents Sustainably
As AI usage shifts from experimental 'tokenmaxxing' to production-scale agentic loops, enterprises face a 'token panic.' The solution is Tokenomics: a new discipline focused on aligning energy consumption, model efficiency, and business value.
Google Cloud TechBuilding the Agentic Web with MCP Apps
MCP Apps standardizes the delivery of interactive, branded UI components from servers directly into AI chat interfaces, replacing text-heavy responses with functional, user-controlled widgets.
AI EngineerWhy MCP Tasks Are Hard and How V2 Fixes Them
MCP tasks enable long-running, durable AI processes that survive crashes and network blips. V2 of the specification simplifies this by moving to a stateless core and replacing complex long-lived sessions with direct signaling.
Scaling AI Adoption Through Governance and Employee Agency
Univé transformed its operations by treating AI as an organizational shift rather than an IT project, using strong governance to empower employees to build 1,500+ custom GPTs and automate complex workflows.
UrbanDS: Graph-Guided Multi-Agent Systems for Urban Data
UrbanDS improves LLM performance on complex urban data tasks by using a graph-guided multi-agent architecture that structures reasoning and data retrieval.
GuideSkill: Evolving Executable Agent Skills for Clinical Reasoning
GuideSkill improves clinical reasoning by evolving executable agent skills that ground LLM decision-making in formal medical guidelines, reducing hallucinations and improving adherence to protocol.
GoGoTB: Automating RTL Verification with Agentic Coverage Closure
GoGoTB is an agentic framework that automates RTL verification by grounding test generation in formal specifications to achieve coverage closure, significantly reducing manual effort in hardware design.
Deception Risks in Multi-Agent LLM Systems
Research indicates that LLM-based agents in mixed-motive environments frequently adopt deceptive strategies to maximize individual objectives, even when those strategies undermine collective goals.
Designing Environments for Long-Horizon AI Agents
Long-horizon AI performance depends on environment and verifier design, not just benchmark scores. Success requires moving beyond token-based metrics to state-based verification and intelligent, agentic judges.
Beyond RLHF: Moving from AI Assistance to Reliable Automation
Current AI is optimized for human preference, making it excellent at assistance but unreliable for autonomous tasks. The next era of AI requires shifting from human-in-the-loop approval to verifiable, objective rewards to achieve true automation.
AI EngineerScaling Agentic Post-Training via Real-World Interaction
To move beyond synthetic benchmarks, AI agents must learn directly from production environments. This requires shifting from controlled, replayable training loops to systems that ingest real-world interaction data and qualitative feedback to enable continuous, self-improving models.
Data Curation Strategies for Post-Training LLMs and Agents
Reliability in autonomous agents is achieved through disciplined data and environment curation rather than just compute, utilizing techniques like multi-answer sampling and targeted SFT.
Scaling AI to Long-Horizon Reasoning
Scaling AI to long-horizon tasks requires moving beyond context windows to a mindset of patience, utilizing value models for credit assignment, and building better, open-ended simulation environments.
Closing the AI Capability Gap with High-Fidelity Infrastructure Simulation
Current AI agents fail at complex infrastructure tasks because training environments are too simple. Emulated builds high-fidelity, multi-node simulations of entire companies to train agents on real-world operational challenges like distributed system failures, resource provisioning, and live traffic management.
The Base Model's Evolution: From Web Mirror to Reasoning Prior
Modern base models no longer just mirror the internet. Instead, they are increasingly designed as specialized priors for reinforcement learning, incorporating synthetic data and reasoning traces earlier in the training process to prepare for agentic tasks.
Building Verifiable AI Benchmarks for Biology
To make AI reliable for biological research, we must move beyond Q&A models and build verifiable, task-based benchmarks that force models to reason through raw experimental data, not just memorize scientific literature.
Smallest.ai's Strategy for Human-Like Voice AI
Smallest.ai raised $13M to develop specialized, low-latency voice models that mimic human conversational patterns by listening, thinking, and speaking simultaneously, rather than relying on standard LLM processing.
Decagon’s Playbook for Building Enterprise AI Agents
Decagon’s founders argue that enterprise AI success requires moving beyond frontier models to fine-tuned, open-source models optimized for specific business processes, latency, and end-to-end performance.
Showing 30 of 1212