Homus

AI Engineer World's Fair 2026: Online Track

70 talks · Online Track · by AI Engineer

Online Track

Build Systems, Not Code

Angie Jones · Agentic AI Foundation · 20 min

Thiết kế agent vẫn là software engineering: mười nguyên tắc từ systems thinking tới maintainability, minh hoạ qua agent tìm nhà Relocation Scout.

Production Evals For Agentic AI Systems

Nishant Gupta · Meta · 8 min

Eval cho agent phải đo hành vi cả hệ thống: scenario offline, production telemetry, human review, drift, trace và metric reliability gắn với kết quả kinh doanh.

Recursive Coding Agents

Raymond Weitekamp · OpenProse · 24 min

Áp nguyên lý recursive language model (RLM) vào coding agent: rubric RLM, Y-Pi, Claude Code dynamic workflows và OpenProse để agent đáng tin cậy hơn.

The Log Is The Agent

Ishaan Sehgal · Omnara · 15 min

Agent không phải model hay runtime mà là append-only log: từ đó có reliability, scaling, forking, migration, và ai giữ log là người sở hữu agent.

The Miranda Hypothesis: How Hamilton (the Musical) Poisoned Your Persona Evals

Jacob E. Thomas · Results Generation · 58 min

Vì sao persona eval chấm fluency không bắt được persona ghép lẫn văn hoá sai thời đại, và instrument pre-registered với nhà sử học để đo fidelity.

A Genius With Amnesia

Victor Savkin · Nx · 20 min

Agent bị giới hạn trong một repo và quên mọi session cũ; Polygraph gỡ cả hai bằng synthetic monorepo và trace dùng chung giữa các agent.

Agents in Production: How OpenGov Built and Scaled OG Assist

Gabe De Mesa · OpenGov · 19 min

OpenGov chạy OG Assist trong production: agent loop Effect-native thay LangGraph, A2A, evals, human approval, sandbox, rolling summary và tracing.

Stop Writing Tone Instructions. Layer Them.

Isadora Martin-Dye · Isadora & Co · 21 min

Tách brand voice thành bốn layer: identity bất biến, mode theo tình huống, voice neo vào ví dụ, và một veto deterministic chặn output sai trước khi tới khách.

Turn 10,994 Notes Into Your Agents' Memory

Paul Iusztin & Louis-François Bouchard · Decoding AI & Towards AI · 40 min

Biến hơn 10.000 note thành memory cho agent bằng plain file: deep research trên second brain, ba layer raw, index.yaml, wiki, không cần vector DB.

Building an Autonomous Engineering Org

Angie Jones · Agentic AI Foundation · 18 min

Cách Block đưa 3.500 engineer từ dùng AI trong IDE lên stage 5 autonomous: AI Champions, AI-friendly repo, Builder Bot, world model, và cái giá về con người.

Agents Building Agents

Alfonso Graziano · Nearform · 30 min

Dùng coding agent để tự cải thiện AI agent: vòng lặp hypothesis trên Golden dataset (18% lên 83%) và phân tích trace có feedback từ user thật.

AI-Driven Multi-Document Correlation for Enterprise Financial Compliance and Fraud Detection

Varsha Shah · Tata Consultancy Services · 19 min

Gian lận nằm giữa các tài liệu: nối payroll, thuế, mua hàng bằng graph, chấm risk score theo xác suất, chuẩn hoá giữa các jurisdiction.

AI System Design: From Idea to Production

Apoorva Joshi · MongoDB · 29 min

Khung bốn phase thiết kế hệ thống AI từ requirement, kiến trúc, eval tới production, áp lên một hệ thống duyệt hồ sơ bảo hiểm y tế.

Browser Agents Don't Need Better Models. They Need Better Eyes.

Kushan Raj · Sarvam AI · 4 min

Browser agent chậm và hay kẹt vì nhìn trang quá tệ: một bản markdown cả trang (~1.800 token) cộng feedback sau mỗi click giúp model rẻ chạy nhanh và đúng.

Bypassing the Multimodal Tax: Framework-Free Hybrid RAG, Raw SQL RRF, and Live UI Telemetry

Abed Matini · Ogilvy · 46 min

Chatbot FAQ chạy local: Docling ra Markdown, bốn chunking strategy, hybrid search trong Postgres gộp bằng RRF, guardrail bằng code và trace với Langfuse.

HTML is All You Need (for Agents to Make Graphics)

Amol Kapoor · Nori Agentic · 7 min

Vì sao agent vẽ dở trên canvas và SVG nhưng làm slide, docs, cả video rất tốt khi viết bằng HTML: đổi medium, đừng đổi model.

OpenClaw in Your Hand: Building a Physical AI Terminal for Local LLM Agents

Lech Kalinowski · Callstack · 25 min

Build Vault: terminal cầm tay hai màn hình OLED và e-paper trên ESP32, điều khiển agent OpenClaw với LLM local và chơi RPG text do LLM dẫn dắt.

Research to Reality: Bringing Frontier ML Research to Production

Vaidas Razgaitis · Higharc · 13 min

Ba đòn bẩy để đưa prototype ML vào production: tài liệu bàn giao RPT, monorepo microservice tách rời, và phân rã prototype thành stacked PR.

Structuring the Unstructured: Advanced Document Parsing for AI Workflows

Cedric Clyburn · Red Hat · 21 min

Biến PDF, bảng, hình ảnh thành Markdown/JSON cho RAG và agent bằng Docling: chạy local trên CPU, chunkless RAG, Docling Serve và MCP server.

The 100-Tool Agent Is a Trap: Scaling with Semantic Routers and JIT Context

Sohail Shaikh & Ankush Rastogi · Prosodica · 28 min

Nạp mọi tool vào prompt làm accuracy rơi từ 78% xuống 13% ở 741 tool; semantic routing (RAG cho tool) với JIT context giữ trên 83% và cắt 99% token.

User Signal Dies at the Retrieval Boundary

Sonam Pankaj · StarlightSearch · 16 min

Vì sao tín hiệu eval chết trong dashboard, và cách dùng utility score để re-rank memory theo outcome, giúp agent tự cải tiến lúc runtime.

Voice In, Visuals Out: The Agony and the Ecstasy

Allen Pike · Forestwalk Labs · 13 min

Vì sao voice-in, visuals-out là UX tốt nhất cho AI, và ba trụ cột giữ phản hồi dưới một giây: model nhanh, inference chu kỳ ngắn, prefix caching ổn định.

We Cut 94% of Our AI Coding Tokens With a Local Code Index. Here's the Architecture.

Rajkumar Sakthivel · Tesco · 11 min

90% chi phí AI coding là input: local code index với Tree-sitter, hybrid search và confidence score đơn giản cắt 94% token gửi lên model.

Building Great Agent Skills: The Missing Manual

Matt Pocock · AI Hero · 21 min

Checklist bốn bước để viết agent skill tốt: chọn trigger, chia steps và reference, lái agent bằng leading words, và tỉa skill bằng deletion test.

Your Agent Failed in Prod. Good Luck Reproducing It.

Tisha Chawla & Susheem Koul · Microsoft · 14 min

Vì sao temperature 0 không làm agent deterministic, và cách record ở boundary để replay một lỗi prod, stub LLM rồi biến trace thành test case.

Building Deterministic Infrastructure for Non-Deterministic AI Agents

Nishant Gupta · Meta · 7 min

Model là stochastic nhưng infrastructure phải deterministic: retry storm, model chỉ đề xuất, agent control plane, trace, memory consistency và safety nhiều lớp.

Frontier results, on device

RL Nabors · Arize · 31 min

Thay lời gọi frontier model bằng SLM chạy local: golden dataset, capability eval với Phoenix, chọn SAGE model, prompt few-shot và post-processing.

The Agentic AI Engineer

Benedikt Sanftl & Burak Cemil Özafşar · Mutagent · 35 min

Chạy vòng đời agent như một loop agentic: spec, build, eval-driven development, diagnose trace production thành eval mới, và demo diagnostics agent.

The Future Is Domain-Specific Agents

Justin Schroeder · StandardAgents · 31 min

Composition over inheritance cho agent: thay vì nhồi MCP và skill vào một agent lớn, ghép nhiều agent nhỏ theo domain, tiết kiệm token, dùng được small model.

The Prompt is the Platform

Dominik Tornow · Resonate HQ · 18 min

Khi agent sinh được implementation, sản phẩm là specification: Resonate dùng deterministic simulation để agent tự design rồi build Resonate trên NATS.

Using RL-based Agent to Detect and Remediate ETL Pipeline Failures

Anna Marie Benzon · University of the Philippines Diliman · 15 min

Agent tự xử lý lỗi ETL trên AWS: rules xác lập facts, Q-learning chọn action có giới hạn, safety layer bên ngoài giữ quyền escalate; MTTR từ ngày xuống phút.

You Can't Prompt the Room: The Last Skill AI Won't Replace

Balázs Horváth · VisualLabs · 16 min

Khi build đã rẻ, phần đắt là quyết định build gì: story mapping, user story, bốn câu hỏi về value và lối tư duy VAD trước khi giao việc cho agent.

The Prompt Is Still a Punch Card

Ted Johnson · JoinIn AI · 20 min

Prompt vẫn là protocol batch của punch card: channel, expression, protocol, và vì sao giao diện AI phải tham gia vào hội thoại thay vì chờ Submit.

Continual Learning for AI Agents: From Failures to Durable Improvements

Soheil Feizi · RELAI · 22 min

Biến log và feedback production thành learning environment replay được, sửa agent ở đúng layer (model, harness, memory) mà không gây regression.

MCP Apps: Primitives, Discovery, and the Future of Software

Pietro Zullo · Manufact · 29 min

MCP Apps: tool trả UI vào chat, các primitive setState, ui/message, streaming input, giấu dữ liệu khỏi model, và cách submit lên store của ChatGPT, Claude, Cursor.

The Missing Layer After Launch

Raphael Kalandadze · Wandero AI · 19 min

Sau khi ship agent: bốn operating agent (log-monitor, PR-review, session-analyzer, QA computer-use) giúp khép vòng loop, đo sức khoẻ và sửa lỗi production.

Your AI Product Will Fail Unless You Can Explain It

Veronica Hylak · Hey AI · 6 min

Ba bước biến AI product phức tạp thành câu chuyện ai cũng hiểu trong một chuyến thang máy: chỉ ra vết thương, làm product click, cho thấy before và after.

SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale

Rishi Desai · Abundant AI · 13 min

Benchmark 20 task cỡ project cho coding agent: agent tốt nhất chỉ đạt 26%, và vì sao verifier đa kênh, CUA, anti-cheat là nút thắt thật.

Build AI Systems for Discernment, Not Approval

Angel Ortmann Lee · Duolingo · 26 min

Duolingo thấy proctor chấp nhận 50% cờ gian lận giả; sửa interaction chứ không sửa model, và các nguyên tắc thiết kế để con người phân định thay vì bấm duyệt.

500 people vibe-coded for 30 days. I was one of them.

Sanja Grbic · Automattic · 18 min

Radical Speed Month ở Automattic: 501 người, 794 project trong 30 ngày, và ba project biến một product designer thành design engineer nhờ AI.

Beyond the Harness: A Journey Towards Adaptive Engineering

Rajiv Chandegra · Annicha Labs · 37 min

Fixed harness hợp với bài toán complicated, còn thế giới thật là complex: adaptive engineering để harness tự nảy sinh từ tương tác giữa các agent.

GTM Is You

Victoria Melnikova · Evil Martians · 13 min

Bottleneck của dev tool năm 2026 là distribution: PMF Compass, sáu bước GTM hygiene và bài học từ founder SF cho thấy personal brand là moat.

How we taught agents to use good retrieval

Hanna Lichtenberg & Aamir Shakir · Mixedbread · 14 min

Vì sao agent viết query keyword vô nghĩa, và cách Mixedbread dùng harness bốn search tool cùng SFT và RL để dạy agent dùng semantic search đúng cách.

Respect The Process

Andrew Dumit · Watershed · 17 min

Cách Watershed để coding agent tự do viết code nhưng buộc mọi edit đi qua typed SDK và deterministic execution, giữ process valid, traceable, replayable.

The Pipeline Is Dead

Iris ten Teije · Sky Valley Ambient Computing · 20 min

Vì sao mô hình một artifact đóng băng cho mọi user đang hết lý do tồn tại, và kiến trúc stem cộng divergence riêng cho từng user giải các phần khó ra sao.

What if the harness mattered more than the model?

Aditya Bhargava · Etsy · 32 min

Cùng model, cùng task, chỉ nâng harness qua bảy nấc bằng ngôn ngữ Agency: tool, handler, PFA, feedback loop, subagent và self-optimization với GEPA.

I Run a Fleet of AI Agents Across Three Machines. Here's What Broke.

Kyle Jaejun Lee · KRAFTON · 9 min

Vận hành đội coding agent trên ba máy: hierarchy CEO→worker, state nằm trong file, reset thay compact, review gateway, năm failure và hướng đi Kubernetes.

Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data

Sachin Kumar · LexisNexis · 14 min

Vì sao eval và behavioral monitor mù trước sleeper agent, và cách bắt backdoor bằng Diff-SAE trên activation delta giữa base và fine-tune.

Chat and citations won't save your vertical AI

Atul Ramachandran · Filed · 15 min

Vertical AI phải thiết kế để giao việc chứ không để tham gia: bốn thành phần delegate, teach, monitor, intervene, và đo WAS thay vì WAU.

Every Solo Agent Builder Eventually Reinvents a Worse Version of CI/CD

Sumaiya Shrabony · University of Colorado Denver · 11 min

Solo builder sẽ tự dựng lại 5 thứ của CI/CD; demo 3 cách agent nói dối (voice drift, claim không nguồn, hook trùng) và gate chặn từng cách.

Stop AI Agent Hallucinations: 5 Techniques + Production Patterns

Elizabeth Fuentes Leone · AWS · 55 min

Năm thay đổi trong code, không phải prompt, để agent bớt hallucinate: lọc tool, GraphRAG, swarm kiểm chéo, hook chặn rule, steering tự sửa; demo Strands.

The Factory That Dreams: 39 AI Agents, No Framework

Rushabh Doshi · Machinecraft · 10 min

Nhà máy 100 người không đội data science build company brain Ira: agent chuyên biệt, memory theo tầng, dream cycle ban đêm, không train model.

A Song of Types and Agents

Roberto Stagi · Ratel · 14 min

Vì sao TypeScript vượt Python trên GitHub và đang chiếm application layer của AI agent: coding agents, npm, một codebase, Zod end-to-end.

Semantic Blindness: 500,000 Sensors Confused an LLM

Raahul Singh & Vanč Levstik · Phaidra · 16 min

Vì sao LLM gãy khi phải đọc 500 nghìn tên thiết bị, và cách Phaidra để LLM chỉ lập plan còn code tra cây: 100% chính xác, ít token hơn 300 lần.

The AI bugpocalypse is here. Now what?

Jack Cable · Corridor · 20 min

Frontier model tìm lỗ hổng ngày càng giỏi, nhưng lỗ hổng vẫn thuộc các lớp cũ: loại bỏ cả lớp bằng memory safety, đặt guardrails cho AI coding.

What Does Done Even Mean? Agents and Paperclip's Liveness Model

Dotta · Paperclip · 7 min

Coi done là một object chứ không phải checkbox: tách bó claim, cân liveness với assurance, và các cơ chế control plane của Paperclip cho agent.

ReviewDebt: a practical framework for scoring every pull request

Sachin Gupta · eBay · 25 min

ReviewDebt chấm mỗi PR bằng năm tín hiệu deterministic để đo khoảng trống giữa code coding agent tạo ra và code con người thật sự review.

The UX of AI: Making AI-Powered Apps Your Users Don't Hate

Kathryn Grayson Nanz · Progress Software · 36 min

Năm trụ cột UX cho tính năng AI: trust, clarity, control, transparency, meaningful benefit, kèm pattern thật như citation, action plan, nút stop, undo.

Your Agents Need a Save Button

Hamza Tahir · ZenML · 17 min

Checkpoint state của agent trong một durable runtime để replay run production với model hay tool khác, diff kết quả và quyết định trên cả cohort.

Agent Output Is Not UX: Rendering Layer Your LLM Pipeline Is Missing

Bala Ramdoss · Amazon · 14 min

Lớp generative UI giữa model và màn hình: rendering contract có version, streaming vào typed component, và BFF giữ cho mobile client an toàn.

Can Oncology Workflows Run Without Human Touch?

Anant Shankhdhar · RISA Labs · 17 min

Bốn agent tự động hoá prior authorization cho thuốc ung thư: deterministic check trước, bằng chứng nhiều nguồn để có confidence, reasoning layer cho ca khó.

Build the AI GTM Agent That Knows the Buyer Before the First Message

Dr. Sajjan Kanukolanu · Position² · 26 min

Kiến trúc ba lớp signals, buyer intelligence, action và context graph để AI GTM agent biết người mua trước tin nhắn đầu, kèm bốn chỗ hệ thống gãy.

Don't Let the LLM Drive

Ornella Bahidika & Joel Allou · Microsoft · 7 min

Voice tutor Ace giữ control flow trong harness: bài học là state machine, model chỉ nhận contract hẹp, nhờ đó chạy reliable trên Haiku 4.5.

Enterprise Agents Have a Structure Problem

Ishita Daga · Tesla · 12 min

Data agent doanh nghiệp sai vì thiếu structure chứ không thiếu context: thứ bậc source of truth, context lifecycle có eval, và bài toán preference còn mở.

Medic for Apache Spark: First Aid for Failing Jobs

Drasko Profirovic · Pinterest · 11 min

Pinterest xây agent chẩn đoán Spark job fail: MCP, E2E test harness record/playback, lọc exception, metrics thành hình, rồi multi-agent trên deepagents.

Skills are the New SDKs

Elvin Aghammadzada · DataRobot · 27 min

Vì sao skill là lớp experience mới cho agent: context rot, progressive disclosure, skill vs MCP, cấu trúc SKILL.md và rủi ro của hệ sinh thái skill.

Designing Voice Agents for Real Conversations

Chintan Agrawal & Daniel Wirjo · AWS · 33 min

Turn-taking cho voice agent qua ba level: Silero VAD, turn detection trong STT, rồi VAD cộng Smart Turn; kèm latency budget, LLM TTFT và demo Pipecat.

When Agents Meet Physical Data: The Other Physics of Agent Harnesses

Dmitry Petrov · DataChain · 28 min

Vì sao coding agent hỏng với video và sensor data, và cách dựng data harness bốn phần: see, run, verify, remember, với demo DataChain và Claude Code.

Why Your Agent Disagrees With Itself (And What To Do About It)

Diane Lin · Datadog · 26 min

Agent flip-flop là dấu hiệu của gray zone: dùng disagreement để chọn ca cho người review, rồi thêm semantic và episodic memory thay vì fine-tune.

Your Voice Agent Doesn't Need a Frontier Model

Ornella Bahidika & Joel Allou · Microsoft · 6 min

Voice tutor Ace chọn model nhỏ vì latency budget khoảng 950 ms: state machine và mastery tracking nằm trong code, model chỉ nói, nhờ đó Haiku 4.5 thay được Opus 4.7.