Topics / Evals
Evals
6 talks
Production Evals For Agentic AI Systems
Nishant Gupta · Meta · 8 min
Eval cho agent phải đo hành vi cả hệ thống: scenario offline, production telemetry, human review, drift, trace và metric reliability gắn với kết quả kinh doanh.
The Miranda Hypothesis: How Hamilton (the Musical) Poisoned Your Persona Evals
Jacob E. Thomas · Results Generation · 58 min
Vì sao persona eval chấm fluency không bắt được persona ghép lẫn văn hoá sai thời đại, và instrument pre-registered với nhà sử học để đo fidelity.
AI-Driven Multi-Document Correlation for Enterprise Financial Compliance and Fraud Detection
Varsha Shah · Tata Consultancy Services · 19 min
Gian lận nằm giữa các tài liệu: nối payroll, thuế, mua hàng bằng graph, chấm risk score theo xác suất, chuẩn hoá giữa các jurisdiction.
User Signal Dies at the Retrieval Boundary
Sonam Pankaj · StarlightSearch · 16 min
Vì sao tín hiệu eval chết trong dashboard, và cách dùng utility score để re-rank memory theo outcome, giúp agent tự cải tiến lúc runtime.
Continual Learning for AI Agents: From Failures to Durable Improvements
Soheil Feizi · RELAI · 22 min
Biến log và feedback production thành learning environment replay được, sửa agent ở đúng layer (model, harness, memory) mà không gây regression.
Why Your Agent Disagrees With Itself (And What To Do About It)
Diane Lin · Datadog · 26 min
Agent flip-flop là dấu hiệu của gray zone: dùng disagreement để chọn ca cho người review, rồi thêm semantic và episodic memory thay vì fine-tune.