name: calendar-scheduling-eval description: Tests an agent's capacity to handle complex constraint satisfaction by managing and resolving conflicting schedules derived purely from raw natural language descriptions. It simulates personal assistant scenarios where the model must maintain a high-density representation of events to detect conflicts accurately. Use when the user wants to benchmark on Calendar Scheduling, or asks about evaluating this task. Reports Exact Match (EM). metadata: skill_kind: dataset_eval source_arxiv: 2604.14930 bibtex_key: lv2026ieascache confidence: high
calendar-scheduling-eval
IE as Cache: Information Extraction Enhanced Agentic Reasoning — Hang Lv et al. (2026) (arXiv:2604.14930, 2026)
What this evaluates
Tests an agent's capacity to handle complex constraint satisfaction by managing and resolving conflicting schedules derived purely from raw natural language descriptions. It simulates personal assistant scenarios where the model must maintain a high-density representation of events to detect conflicts accurately.
Datasets
- Calendar Scheduling — total ?; splits: test (-1)
Metrics
Exact Match (EM)(primary) — range: [0, 1]- 1 if the predicted schedule exactly satisfies all constraints and matches the ground-truth solution, 0 otherwise.
Input / output format
Input: Raw natural language descriptions of scheduling constraints and events.
Output: A final schedule or decision satisfying the target constraints.
Scoring recipe
def score_em(pred, gold):
return 1.0 if pred.strip().lower() == gold.strip().lower() else 0.0
# Average over test set
em_scores = [score_em(p, g) for p, g in zip(predictions, gold_answers)]
final_metric = sum(em_scores) / len(em_scores)
Common pitfalls
- Frequent context switching in agentic baselines (e.g., ReAct) can fragment constraint tracking and introduce noise.
- Smaller models struggle significantly with multiple simultaneous constraints, often requiring explicit cache mechanisms to maintain performance.
- Zero-shot extraction scenarios may be sensitive to prompt structure, causing minor baselines to occasionally outperform more complex methods.
Evidence (verbatim from paper)
For TACT and Calendar Scheduling, we report Exact Match (EM), which directly measures whether the produced answer satisfies the target decision/constraints.
Citation
@misc{lv2026ieascache,
title={IE as Cache: Information Extraction Enhanced Agentic Reasoning},
author={Hang Lv et al. (2026)},
year={2026},
note={arXiv:2604.14930}
}
- arXiv: 2604.14930
TDD Red-Green-Refactor
Testing
Skill that guides Claude through the complete TDD cycle.
Web Accessibility Audit
Testing
Performs a comprehensive web accessibility audit following WCAG standards.
UAT Test Case Generator
Testing
Generates structured and comprehensive user acceptance test cases.