Homus

Topics / Evals

Evals

6 talks

Production Evals For Agentic AI Systems

Nishant Gupta · Meta · 8 min

Eval cho agent phải đo hành vi cả hệ thống: scenario offline, production telemetry, human review, drift, trace và metric reliability gắn với kết quả kinh doanh.

The Miranda Hypothesis: How Hamilton (the Musical) Poisoned Your Persona Evals

Jacob E. Thomas · Results Generation · 58 min

Vì sao persona eval chấm fluency không bắt được persona ghép lẫn văn hoá sai thời đại, và instrument pre-registered với nhà sử học để đo fidelity.

AI-Driven Multi-Document Correlation for Enterprise Financial Compliance and Fraud Detection

Varsha Shah · Tata Consultancy Services · 19 min

Gian lận nằm giữa các tài liệu: nối payroll, thuế, mua hàng bằng graph, chấm risk score theo xác suất, chuẩn hoá giữa các jurisdiction.

User Signal Dies at the Retrieval Boundary

Sonam Pankaj · StarlightSearch · 16 min

Vì sao tín hiệu eval chết trong dashboard, và cách dùng utility score để re-rank memory theo outcome, giúp agent tự cải tiến lúc runtime.

Continual Learning for AI Agents: From Failures to Durable Improvements

Soheil Feizi · RELAI · 22 min

Biến log và feedback production thành learning environment replay được, sửa agent ở đúng layer (model, harness, memory) mà không gây regression.

Why Your Agent Disagrees With Itself (And What To Do About It)

Diane Lin · Datadog · 26 min

Agent flip-flop là dấu hiệu của gray zone: dùng disagreement để chọn ca cho người review, rồi thêm semantic và episodic memory thay vì fine-tune.