# Evals

Topic · 9 talks. Cách đo chất lượng AI: benchmark, rubric, LLM-as-judge, eval gate trong pipeline.

Canonical: https://homus.dev/topics/evals

- [Skill-creator mới: test, đo và tối ưu Claude skill](https://homus.dev/talks/anthropic-just-dropped-claude-code-skills-2-0.md): Ray Amjad, Ray Amjad. Skill-creator mới của Anthropic: hai loại skill, chạy evals và A/B test mù có/không skill, benchmark, tối ưu description để skill trigger ổn định.
- [Tạo skill đầu tiên trong Claude](https://homus.dev/talks/how-to-create-skills-in-claude.md): Tom Nassr, Tom Nassr | XRAY. Tạo skill YouTube chapters bằng skill-creator trên claude.ai, dùng nó trong chat mới, rồi để Cowork chạy test và benchmark để tinh chỉnh skill.
- [Nhân viên Anthropic thật sự dùng Claude skills thế nào](https://homus.dev/talks/how-anthropic-employees-actually-use-claude-skills.md): Austin Marchese, Austin Marchese. Năm bài học từ playbook nội bộ của Anthropic: bốn skill type, scripts và templates, verifier, gotchas và cách viết description để skill tự trigger.
- [Eval trong production cho hệ thống agentic AI](https://homus.dev/talks/production-evals-for-agentic-ai-systems.md): Nishant Gupta, Meta. Eval cho agent phải đo hành vi cả hệ thống: scenario offline, production telemetry, human review, drift, trace và metric reliability gắn với kết quả kinh doanh.
- [Giả thuyết Miranda: vở musical Hamilton đã đầu độc persona eval của bạn ra sao](https://homus.dev/talks/the-miranda-hypothesis-how-hamilton-the-musical-poisoned-your-persona-evals.md): Jacob E. Thomas, Results Generation. Vì sao persona eval chấm fluency không bắt được persona ghép lẫn văn hoá sai thời đại, và instrument pre-registered với nhà sử học để đo fidelity.
- [Đối chiếu nhiều tài liệu bằng AI để kiểm tra tuân thủ tài chính và phát hiện gian lận](https://homus.dev/talks/ai-driven-multi-document-correlation-for-enterprise-financial-compliance-and-fraud-detection.md): Varsha Shah, Tata Consultancy Services. Gian lận nằm giữa các tài liệu: nối payroll, thuế, mua hàng bằng graph, chấm risk score theo xác suất, chuẩn hoá giữa các jurisdiction.
- [Tín hiệu từ user chết ở ranh giới retrieval](https://homus.dev/talks/user-signal-dies-at-the-retrieval-boundary.md): Sonam Pankaj, StarlightSearch. Vì sao tín hiệu eval chết trong dashboard, và cách dùng utility score để re-rank memory theo outcome, giúp agent tự cải tiến lúc runtime.
- [Continual learning cho AI agent: từ thất bại tới cải tiến bền vững](https://homus.dev/talks/continual-learning-for-ai-agents-from-failures-to-durable-improvements.md): Soheil Feizi, RELAI. Biến log và feedback production thành learning environment replay được, sửa agent ở đúng layer (model, harness, memory) mà không gây regression.
- [Vì sao agent bất đồng với chính nó, và nên làm gì](https://homus.dev/talks/why-your-agent-disagrees-with-itself-and-what-to-do-about-it.md): Diane Lin, Datadog. Agent flip-flop là dấu hiệu của gray zone: dùng disagreement để chọn ca cho người review, rồi thêm semantic và episodic memory thay vì fine-tune.

Nguồn: https://homus.dev/topics/evals · cập nhật 2026-10-09
