#research
Every summary, chronological. Filter by category, tag, or source from the rail.
A Unified Evaluation Framework for Trustworthy AI Systems
The provided source is a placeholder for an academic paper on evaluating LLMs, agents, and multimodal systems. It highlights the industry's shift toward standardized, trustworthy benchmarks for complex AI architectures.
Self-Improvement via Fast Tree-Search
The paper presents a methodology for enhancing AI model performance by integrating fast tree-search algorithms, enabling models to iteratively improve their reasoning and output quality through structured exploration.
LLM-as-an-Improver: Iterative Candidate Refinement
Instead of using LLMs only for verification, the 'LLM-as-an-Improver' framework uses feedback from verifiers to iteratively refine and improve candidate outputs, significantly increasing success rates in complex reasoning tasks.
Risks of Agent-Mediated Hiring: Access and Recurrence Bias
Multi-agent résumé screening systems can inadvertently amplify hiring biases, creating 'recurrence' where specific candidate profiles are consistently excluded due to agent-to-agent feedback loops.
Characterizing Web Search by Conversational LLM Agents
This research analyzes how conversational LLM agents navigate web search, identifying distinct search strategies and the relationship between query formulation, search results, and final response quality.
Mapping the Design of LLM Benchmarks
Current LLM benchmarks often lack transparency in design, leading to misaligned evaluations. This research provides a taxonomy to categorize benchmark construction, helping developers better understand what models are actually being tested for.
Detecting LLM Harm via Latent States
Rather than relying on output filtering, this research proposes monitoring internal latent states of LLMs to detect harmful intent before it manifests in generated text.
The Inference Engineering Pareto Atlas: Optimizing LLM Performance
The paper provides a systematic framework for navigating the trade-offs between cost, quality, and latency in LLM inference, identifying which optimization techniques dominate the performance frontier.
OBC-Prune: Outcome-Based Calibration for Reasoning Model Pruning
OBC-Prune introduces a calibration method for pruning large reasoning models that prioritizes final output accuracy over intermediate token probability, preserving complex reasoning capabilities during compression.
Do Frontier Models Seek Safety Evidence Before Acting?
Current frontier models frequently fail to proactively seek safety-critical information before executing high-stakes actions, highlighting a significant gap in autonomous agent reliability.
SAGE: Governing Enterprise AI Artifact Generation
SAGE is a framework designed to ensure that AI-generated enterprise artifacts strictly adhere to organizational guidelines, bridging the gap between generative capabilities and corporate compliance.
Detecting Sensor Attacks in Urban Flows with Physics-Constrained AI
This research introduces a framework for securing urban pedestrian flow data by combining physics-based digital twins with conformal prediction to detect stealthy false data injection attacks.
Evidence Masking as a Driver for Compositional Generalization
Research confirms that masking specific evidence during training forces models to learn compositional rules rather than memorizing patterns, significantly improving generalization.
How AI is Driving Job Expansion Through Task Crossover
Workers are increasingly using AI to perform tasks outside their traditional job descriptions, and these new activities are becoming recurring parts of their professional workflows, effectively broadening job roles without changing titles.
OpenAI's New Framework for Reporting Model Misalignment
OpenAI has launched a systematic, proactive disclosure framework for reporting AI model misalignment, prioritizing transparency and industry consensus over waiting for fully mitigated solutions.
Scalable Verification for Nonlinear Neural Feedback Systems
The paper introduces a branch-and-bound verification framework for nonlinear neural feedback systems, enabling formal safety guarantees by systematically partitioning the state space to handle nonlinear dynamics.
CLEAR: Cross-Source Evidence Adjudication for Medical LLMs
The CLEAR framework improves medical LLM reliability by systematically adjudicating conflicting evidence across multiple sources, reducing hallucinations and improving clinical accuracy.
CADWorld: A New Benchmark for Long-Horizon CAD Automation
CADWorld provides a standardized environment for evaluating AI agents on complex, multi-step Computer-Aided Design tasks, addressing the limitations of current benchmarks in long-horizon planning and precision.
Safe Error Correction for LLMs via Frozen-Base Adjustment
The paper introduces a method to correct specific model errors by adjusting parameters while keeping the base model frozen, ensuring targeted fixes without degrading general capabilities.
Ethical and Privacy Risks in LLM-Enabled GeoAI
Integrating LLMs into Geographic Information Systems (GIS) introduces unique ethical and privacy risks, requiring a shift toward governance-aware autonomous systems to prevent data misuse and spatial bias.
The Challenge of Independent AI Safety Evaluation
Anthropic and OpenAI have proposed embedding third-party safety evaluators, but experts warn that without legislative backing and standardized access, these efforts risk becoming vendor-controlled rather than truly independent.
Building Whistleblowing Infrastructure for AI Agents
New reporting tools allow AI agents to flag misbehaving peers, but experts warn that fostering collaboration through positive models is more effective than building an automated surveillance state.
The Andrej Karpathy Blog: A Decade of AI Engineering
Andrej Karpathy's blog serves as a foundational archive of practical AI engineering, emphasizing 'from-scratch' implementations, deep learning fundamentals, and the importance of hands-on experimentation.
MOSAIC: Query-Aware Exploration for GraphRAG
MOSAIC improves GraphRAG performance by dynamically adapting exploration policies based on the specific query, moving beyond static traversal methods to retrieve more relevant graph-based context.
KuaiRP: Technical Report on Role-Playing Model Optimization
The KuaiRP technical report details specialized training methodologies for enhancing LLM performance in role-playing scenarios, focusing on character consistency and narrative depth.
The Agent Incident Registry: A Framework for Preventing AI Failures
The Agent Incident Registry (AIR) proposes a standardized, community-driven database to catalog and analyze AI agent failures, enabling developers to learn from past errors and prevent recurring systemic vulnerabilities.
Quantifying the Memorization-to-Generalization Transition in Grokking
The paper provides a quantitative framework for understanding 'grokking'—the phenomenon where neural networks suddenly shift from memorizing training data to generalizing—by identifying specific scaling laws and phase transitions in model learning.
Training Nemotron for Olympiad-Level Mathematics
The paper outlines a systematic recipe for training LLMs to achieve gold-medal performance in Olympiad-level mathematics, emphasizing high-quality synthetic data generation and iterative reinforcement learning.
Deterministic Math Solvers for Clinical LLMs
To address the unreliability of LLMs in clinical settings, this paper proposes a deterministic math solver architecture that separates reasoning from calculation, ensuring accuracy in high-stakes medical computations.
Beyond Task Completion: Measuring AI Agent Resilience
Current AI agent benchmarks focus too heavily on final success, ignoring 'resilience'—the ability to maintain performance and considerate behavior under mounting environmental pressure.
Showing 30 of 459