Evaluating Agent Personalization

PersonaTrail addresses a critical gap in current web agent research: the inability to measure how well an agent adapts to individual user preferences. While most benchmarks focus on task completion (e.g., 'book a flight'), they ignore the 'how'—the specific stylistic, privacy, or navigational choices a human user would make. By utilizing 'browsing trails'—sequences of historical user interactions—the benchmark provides a structured way to test if an agent can replicate a specific persona's decision-making process during web navigation.

The Role of Browsing Trails

The core innovation of PersonaTrail is the use of historical browsing data as a ground-truth signal for agent behavior. Instead of evaluating agents solely on whether they reach a destination, the benchmark measures alignment with the user's past patterns. This allows researchers to quantify:

  • Preference Adherence: Does the agent choose the same types of sites or services the user historically prefers?
  • Contextual Consistency: Can the agent maintain a persona's constraints (e.g., budget, accessibility needs, or privacy settings) across a multi-step browsing session?
  • Adaptability: How well does the agent handle new websites while maintaining the established persona profile?

This approach moves the field toward more human-centric AI, where agents act as personal assistants rather than generic task-executors.