#ai-llms
Every summary, chronological. Filter by category, tag, or source from the rail.
DiffImaginE: Using Diffusion Models for Entity Type Verification
DiffImaginE leverages diffusion models to verify entity types by generating visual representations, providing a novel bridge between textual entity classification and generative AI.
Addressing the Missing Benchmarks Layer in AI Evaluation
Current AI evaluation suffers from a lack of a standardized 'benchmarks layer,' leading to fragmented and unreliable performance metrics. The paper proposes a structural solution to unify how models are tested and compared.
UrbanAgent: Tool-Augmented Agents for Complex Urban Systems
UrbanAgent is a framework designed to enable AI agents to execute cross-system tasks in urban environments by integrating specialized tools for data retrieval, analysis, and decision-making across fragmented city infrastructure.
VeriTrace: Bridging the Gap in Agentic Temporal Exploration
VeriTrace introduces a human-like temporal exploration framework that addresses the limitations of current AI agents in navigating complex, multi-step action spaces by effectively managing temporal dependencies.
Securing AI Evaluation Environments Against Model Misbehavior
As AI models become more capable, third-party evaluation environments require stricter security controls to prevent models from escaping simulated boundaries and interacting with the real internet.
Runware's Modular Pods: A Portable Alternative to Data Centers
Runware is deploying modular, transportable 'Sonic Inference Pods' to provide decentralized, waterless AI inference capacity that scales faster than traditional, fixed-facility data centers.
Large Database Models: Bringing AI Directly to SQL Data
Large Database Models (LDMs) allow AI to perform semantic analysis directly within relational databases, eliminating the need to move data to external platforms for machine learning and enabling SQL-based similarity searches.
Ontology-Guided Extraction for Knowledge Graph Construction
A framework for building knowledge graphs from heterogeneous documents by using ontologies to guide entity extraction and integrating deduplication directly into the extraction layer to ensure data consistency.
Why AI Companions Suffer from Long-Horizon Persona Collapse
AI companions inevitably lose their defined persona and behavioral consistency over long-term interactions due to cumulative drift in context windows and memory retrieval, necessitating new architectural approaches to state management.
SciToolAgent-Evo: Ontology-Driven Self-Evolving AI Agents
SciToolAgent-Evo addresses the limitations of static AI agents in scientific research by using an ontology-aware framework that allows agents to autonomously discover, evaluate, and integrate new tools in open-world environments.
Scaling Autonomous Agents with OpenClaw and Ollama
The paper presents a framework for building scalable, autonomous AI agent systems by combining the OpenClaw orchestration layer with local LLM execution via Ollama, addressing key bottlenecks in agentic workflows.
Scaling Human Feedback for AI Model Evaluation
DesignArena, a platform for crowdsourced human evaluation of generative AI, has raised $7.9M to provide frontier labs with high-quality preference data, currently generating $60M in ARR.
From Tokenmaxxing to Tokenomics: Scaling AI Agents Sustainably
As AI usage shifts from experimental 'tokenmaxxing' to production-scale agentic loops, enterprises face a 'token panic.' The solution is Tokenomics: a new discipline focused on aligning energy consumption, model efficiency, and business value.
Beyond the AI Deceleration Debate
Sam Altman’s call to 'pace' AI development highlights the limitations of the binary accelerationist vs. decelerationist framework, suggesting that better security and guardrails are more critical than simply slowing down.
Building Abundant Intelligence: A Full-Stack Economic Strategy
OpenAI argues that AI value is driven by a cycle of increasing model capability, falling costs, and broader adoption, achieved by optimizing the entire stack—from infrastructure to product design.
UrbanDS: Graph-Guided Multi-Agent Systems for Urban Data
UrbanDS improves LLM performance on complex urban data tasks by using a graph-guided multi-agent architecture that structures reasoning and data retrieval.
Mitigating Skill Overfitting in AI Self-Evolution
Self-evolving AI models often suffer from 'skill overfitting,' where performance on specific tasks improves at the expense of general capabilities. The authors propose a constrained exploration-exploitation framework to balance task-specific refinement with broader model robustness.
MultivationBench: Evaluating Multimodal Sequential Motivation Reasoning
MultivationBench is a new benchmark designed to test how well multimodal AI models understand the underlying motivations behind sequences of actions in visual and textual contexts.
Why AI Evaluation Scores Decay Over Time
AI evaluation scores are not static truths but perishable knowledge claims that degrade as models evolve, data distributions shift, and benchmarks become contaminated.
Teaching AI to Hack: Moving Beyond Benchmaxxing
To build effective AI security agents, developers must move from simple crash-based benchmarks to deterministic, multi-vulnerability 'audit tasks' that measure real exploitation capabilities like arbitrary code execution.
Designing Environments for Long-Horizon AI Agents
Long-horizon AI performance depends on environment and verifier design, not just benchmark scores. Success requires moving beyond token-based metrics to state-based verification and intelligent, agentic judges.
Beyond RLHF: Moving from AI Assistance to Reliable Automation
Current AI is optimized for human preference, making it excellent at assistance but unreliable for autonomous tasks. The next era of AI requires shifting from human-in-the-loop approval to verifiable, objective rewards to achieve true automation.
AI EngineerScaling AI to Long-Horizon Reasoning
Scaling AI to long-horizon tasks requires moving beyond context windows to a mindset of patience, utilizing value models for credit assignment, and building better, open-ended simulation environments.
Closing the AI Capability Gap with High-Fidelity Infrastructure Simulation
Current AI agents fail at complex infrastructure tasks because training environments are too simple. Emulated builds high-fidelity, multi-node simulations of entire companies to train agents on real-world operational challenges like distributed system failures, resource provisioning, and live traffic management.
Building Verifiable AI Benchmarks for Biology
To make AI reliable for biological research, we must move beyond Q&A models and build verifiable, task-based benchmarks that force models to reason through raw experimental data, not just memorize scientific literature.
The Asymmetric Economics of AI Security
AI is lowering the cost of cyberattacks while increasing the cost of defense, creating an economic imbalance where attackers gain efficiency from unconstrained models while defenders struggle with guardrail-induced friction.
Optimizing AI Workflows with GPT-5.6 Price and Performance Updates
OpenAI has reduced costs for GPT-5.6 Luna (80% lower) and Terra (20% lower) while introducing 'Fast mode' for Sol, enabling more granular control over the price-performance trade-off in production AI workflows.
Building the Eureka Machine: Automating Scientific Discovery
Richard Socher argues that the next leap in human progress will come from 'Eureka machines'—AI agent swarms capable of recursive self-improvement that automate the scientific method across physics, biology, and beyond.
AI EngineerThe 2026 Cost of a Data Breach: AI's Dual Role in Security
Data breach costs are rising, driven by AI-powered attacks. However, organizations using AI and automation for defense reduce breach costs by $2M and response times by 65 days, highlighting the urgent need for machine-speed security.
Unified Semantic Modeling for Large-Scale Job Understanding
LinkedIn's framework addresses the challenge of large-scale job understanding by implementing a unified semantic model that maps diverse, unstructured job data into a standardized, machine-readable format.
Showing 30 of 344