AI Engineer World's Fair 2026: Online Track
70 talks · Online Track · by AI EngineerOnline Track
Build Systems, Not Code
Angie Jones · Agentic AI Foundation · 20 min
Thiết kế agent vẫn là software engineering: mười nguyên tắc từ systems thinking tới maintainability, minh hoạ qua agent tìm nhà Relocation Scout.
Production Evals For Agentic AI Systems
Nishant Gupta · Meta · 8 min
Eval cho agent phải đo hành vi cả hệ thống: scenario offline, production telemetry, human review, drift, trace và metric reliability gắn với kết quả kinh doanh.
Recursive Coding Agents
Raymond Weitekamp · OpenProse · 24 min
Áp nguyên lý recursive language model (RLM) vào coding agent: rubric RLM, Y-Pi, Claude Code dynamic workflows và OpenProse để agent đáng tin cậy hơn.
The Log Is The Agent
Ishaan Sehgal · Omnara · 15 min
Agent không phải model hay runtime mà là append-only log: từ đó có reliability, scaling, forking, migration, và ai giữ log là người sở hữu agent.
The Miranda Hypothesis: How Hamilton (the Musical) Poisoned Your Persona Evals
Jacob E. Thomas · Results Generation · 58 min
Vì sao persona eval chấm fluency không bắt được persona ghép lẫn văn hoá sai thời đại, và instrument pre-registered với nhà sử học để đo fidelity.
A Genius With Amnesia
Victor Savkin · Nx · 20 min
Agent bị giới hạn trong một repo và quên mọi session cũ; Polygraph gỡ cả hai bằng synthetic monorepo và trace dùng chung giữa các agent.
Agents in Production: How OpenGov Built and Scaled OG Assist
Gabe De Mesa · OpenGov · 19 min
OpenGov chạy OG Assist trong production: agent loop Effect-native thay LangGraph, A2A, evals, human approval, sandbox, rolling summary và tracing.
Stop Writing Tone Instructions. Layer Them.
Isadora Martin-Dye · Isadora & Co · 21 min
Tách brand voice thành bốn layer: identity bất biến, mode theo tình huống, voice neo vào ví dụ, và một veto deterministic chặn output sai trước khi tới khách.
Turn 10,994 Notes Into Your Agents' Memory
Paul Iusztin & Louis-François Bouchard · Decoding AI & Towards AI · 40 min
Biến hơn 10.000 note thành memory cho agent bằng plain file: deep research trên second brain, ba layer raw, index.yaml, wiki, không cần vector DB.
Building an Autonomous Engineering Org
Angie Jones · Agentic AI Foundation · 18 min
Cách Block đưa 3.500 engineer từ dùng AI trong IDE lên stage 5 autonomous: AI Champions, AI-friendly repo, Builder Bot, world model, và cái giá về con người.
Agents Building Agents
Alfonso Graziano · Nearform · 30 min
Dùng coding agent để tự cải thiện AI agent: vòng lặp hypothesis trên Golden dataset (18% lên 83%) và phân tích trace có feedback từ user thật.
AI-Driven Multi-Document Correlation for Enterprise Financial Compliance and Fraud Detection
Varsha Shah · Tata Consultancy Services · 19 min
Gian lận nằm giữa các tài liệu: nối payroll, thuế, mua hàng bằng graph, chấm risk score theo xác suất, chuẩn hoá giữa các jurisdiction.
AI System Design: From Idea to Production
Apoorva Joshi · MongoDB · 29 min
Khung bốn phase thiết kế hệ thống AI từ requirement, kiến trúc, eval tới production, áp lên một hệ thống duyệt hồ sơ bảo hiểm y tế.
Browser Agents Don't Need Better Models. They Need Better Eyes.
Kushan Raj · Sarvam AI · 4 min
Browser agent chậm và hay kẹt vì nhìn trang quá tệ: một bản markdown cả trang (~1.800 token) cộng feedback sau mỗi click giúp model rẻ chạy nhanh và đúng.
Bypassing the Multimodal Tax: Framework-Free Hybrid RAG, Raw SQL RRF, and Live UI Telemetry
Abed Matini · Ogilvy · 46 min
Chatbot FAQ chạy local: Docling ra Markdown, bốn chunking strategy, hybrid search trong Postgres gộp bằng RRF, guardrail bằng code và trace với Langfuse.
HTML is All You Need (for Agents to Make Graphics)
Amol Kapoor · Nori Agentic · 7 min
Vì sao agent vẽ dở trên canvas và SVG nhưng làm slide, docs, cả video rất tốt khi viết bằng HTML: đổi medium, đừng đổi model.
OpenClaw in Your Hand: Building a Physical AI Terminal for Local LLM Agents
Lech Kalinowski · Callstack · 25 min
Build Vault: terminal cầm tay hai màn hình OLED và e-paper trên ESP32, điều khiển agent OpenClaw với LLM local và chơi RPG text do LLM dẫn dắt.
Research to Reality: Bringing Frontier ML Research to Production
Vaidas Razgaitis · Higharc · 13 min
Ba đòn bẩy để đưa prototype ML vào production: tài liệu bàn giao RPT, monorepo microservice tách rời, và phân rã prototype thành stacked PR.
Structuring the Unstructured: Advanced Document Parsing for AI Workflows
Cedric Clyburn · Red Hat · 21 min
Biến PDF, bảng, hình ảnh thành Markdown/JSON cho RAG và agent bằng Docling: chạy local trên CPU, chunkless RAG, Docling Serve và MCP server.
The 100-Tool Agent Is a Trap: Scaling with Semantic Routers and JIT Context
Sohail Shaikh & Ankush Rastogi · Prosodica · 28 min
Nạp mọi tool vào prompt làm accuracy rơi từ 78% xuống 13% ở 741 tool; semantic routing (RAG cho tool) với JIT context giữ trên 83% và cắt 99% token.
User Signal Dies at the Retrieval Boundary
Sonam Pankaj · StarlightSearch · 16 min
Vì sao tín hiệu eval chết trong dashboard, và cách dùng utility score để re-rank memory theo outcome, giúp agent tự cải tiến lúc runtime.
Voice In, Visuals Out: The Agony and the Ecstasy
Allen Pike · Forestwalk Labs · 13 min
Vì sao voice-in, visuals-out là UX tốt nhất cho AI, và ba trụ cột giữ phản hồi dưới một giây: model nhanh, inference chu kỳ ngắn, prefix caching ổn định.
We Cut 94% of Our AI Coding Tokens With a Local Code Index. Here's the Architecture.
Rajkumar Sakthivel · Tesco · 11 min
90% chi phí AI coding là input: local code index với Tree-sitter, hybrid search và confidence score đơn giản cắt 94% token gửi lên model.
Building Great Agent Skills: The Missing Manual
Matt Pocock · AI Hero · 21 min
Checklist bốn bước để viết agent skill tốt: chọn trigger, chia steps và reference, lái agent bằng leading words, và tỉa skill bằng deletion test.
Your Agent Failed in Prod. Good Luck Reproducing It.
Tisha Chawla & Susheem Koul · Microsoft · 14 min
Vì sao temperature 0 không làm agent deterministic, và cách record ở boundary để replay một lỗi prod, stub LLM rồi biến trace thành test case.
Building Deterministic Infrastructure for Non-Deterministic AI Agents
Nishant Gupta · Meta · 7 min
Model là stochastic nhưng infrastructure phải deterministic: retry storm, model chỉ đề xuất, agent control plane, trace, memory consistency và safety nhiều lớp.
Frontier results, on device
RL Nabors · Arize · 31 min
Thay lời gọi frontier model bằng SLM chạy local: golden dataset, capability eval với Phoenix, chọn SAGE model, prompt few-shot và post-processing.
The Agentic AI Engineer
Benedikt Sanftl & Burak Cemil Özafşar · Mutagent · 35 min
Chạy vòng đời agent như một loop agentic: spec, build, eval-driven development, diagnose trace production thành eval mới, và demo diagnostics agent.
The Future Is Domain-Specific Agents
Justin Schroeder · StandardAgents · 31 min
Composition over inheritance cho agent: thay vì nhồi MCP và skill vào một agent lớn, ghép nhiều agent nhỏ theo domain, tiết kiệm token, dùng được small model.
The Prompt is the Platform
Dominik Tornow · Resonate HQ · 18 min
Khi agent sinh được implementation, sản phẩm là specification: Resonate dùng deterministic simulation để agent tự design rồi build Resonate trên NATS.
Using RL-based Agent to Detect and Remediate ETL Pipeline Failures
Anna Marie Benzon · University of the Philippines Diliman · 15 min
Agent tự xử lý lỗi ETL trên AWS: rules xác lập facts, Q-learning chọn action có giới hạn, safety layer bên ngoài giữ quyền escalate; MTTR từ ngày xuống phút.
You Can't Prompt the Room: The Last Skill AI Won't Replace
Balázs Horváth · VisualLabs · 16 min
Khi build đã rẻ, phần đắt là quyết định build gì: story mapping, user story, bốn câu hỏi về value và lối tư duy VAD trước khi giao việc cho agent.
The Prompt Is Still a Punch Card
Ted Johnson · JoinIn AI · 20 min
Prompt vẫn là protocol batch của punch card: channel, expression, protocol, và vì sao giao diện AI phải tham gia vào hội thoại thay vì chờ Submit.
Continual Learning for AI Agents: From Failures to Durable Improvements
Soheil Feizi · RELAI · 22 min
Biến log và feedback production thành learning environment replay được, sửa agent ở đúng layer (model, harness, memory) mà không gây regression.
MCP Apps: Primitives, Discovery, and the Future of Software
Pietro Zullo · Manufact · 29 min
MCP Apps: tool trả UI vào chat, các primitive setState, ui/message, streaming input, giấu dữ liệu khỏi model, và cách submit lên store của ChatGPT, Claude, Cursor.
The Missing Layer After Launch
Raphael Kalandadze · Wandero AI · 19 min
Sau khi ship agent: bốn operating agent (log-monitor, PR-review, session-analyzer, QA computer-use) giúp khép vòng loop, đo sức khoẻ và sửa lỗi production.
Your AI Product Will Fail Unless You Can Explain It
Veronica Hylak · Hey AI · 6 min
Ba bước biến AI product phức tạp thành câu chuyện ai cũng hiểu trong một chuyến thang máy: chỉ ra vết thương, làm product click, cho thấy before và after.
SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale
Rishi Desai · Abundant AI · 13 min
Benchmark 20 task cỡ project cho coding agent: agent tốt nhất chỉ đạt 26%, và vì sao verifier đa kênh, CUA, anti-cheat là nút thắt thật.
Build AI Systems for Discernment, Not Approval
Angel Ortmann Lee · Duolingo · 26 min
Duolingo thấy proctor chấp nhận 50% cờ gian lận giả; sửa interaction chứ không sửa model, và các nguyên tắc thiết kế để con người phân định thay vì bấm duyệt.
500 people vibe-coded for 30 days. I was one of them.
Sanja Grbic · Automattic · 18 min
Radical Speed Month ở Automattic: 501 người, 794 project trong 30 ngày, và ba project biến một product designer thành design engineer nhờ AI.
Beyond the Harness: A Journey Towards Adaptive Engineering
Rajiv Chandegra · Annicha Labs · 37 min
Fixed harness hợp với bài toán complicated, còn thế giới thật là complex: adaptive engineering để harness tự nảy sinh từ tương tác giữa các agent.
GTM Is You
Victoria Melnikova · Evil Martians · 13 min
Bottleneck của dev tool năm 2026 là distribution: PMF Compass, sáu bước GTM hygiene và bài học từ founder SF cho thấy personal brand là moat.
How we taught agents to use good retrieval
Hanna Lichtenberg & Aamir Shakir · Mixedbread · 14 min
Vì sao agent viết query keyword vô nghĩa, và cách Mixedbread dùng harness bốn search tool cùng SFT và RL để dạy agent dùng semantic search đúng cách.
Respect The Process
Andrew Dumit · Watershed · 17 min
Cách Watershed để coding agent tự do viết code nhưng buộc mọi edit đi qua typed SDK và deterministic execution, giữ process valid, traceable, replayable.
The Pipeline Is Dead
Iris ten Teije · Sky Valley Ambient Computing · 20 min
Vì sao mô hình một artifact đóng băng cho mọi user đang hết lý do tồn tại, và kiến trúc stem cộng divergence riêng cho từng user giải các phần khó ra sao.
What if the harness mattered more than the model?
Aditya Bhargava · Etsy · 32 min
Cùng model, cùng task, chỉ nâng harness qua bảy nấc bằng ngôn ngữ Agency: tool, handler, PFA, feedback loop, subagent và self-optimization với GEPA.
I Run a Fleet of AI Agents Across Three Machines. Here's What Broke.
Kyle Jaejun Lee · KRAFTON · 9 min
Vận hành đội coding agent trên ba máy: hierarchy CEO→worker, state nằm trong file, reset thay compact, review gateway, năm failure và hướng đi Kubernetes.
Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data
Sachin Kumar · LexisNexis · 14 min
Vì sao eval và behavioral monitor mù trước sleeper agent, và cách bắt backdoor bằng Diff-SAE trên activation delta giữa base và fine-tune.
Chat and citations won't save your vertical AI
Atul Ramachandran · Filed · 15 min
Vertical AI phải thiết kế để giao việc chứ không để tham gia: bốn thành phần delegate, teach, monitor, intervene, và đo WAS thay vì WAU.
Every Solo Agent Builder Eventually Reinvents a Worse Version of CI/CD
Sumaiya Shrabony · University of Colorado Denver · 11 min
Solo builder sẽ tự dựng lại 5 thứ của CI/CD; demo 3 cách agent nói dối (voice drift, claim không nguồn, hook trùng) và gate chặn từng cách.
Stop AI Agent Hallucinations: 5 Techniques + Production Patterns
Elizabeth Fuentes Leone · AWS · 55 min
Năm thay đổi trong code, không phải prompt, để agent bớt hallucinate: lọc tool, GraphRAG, swarm kiểm chéo, hook chặn rule, steering tự sửa; demo Strands.
The Factory That Dreams: 39 AI Agents, No Framework
Rushabh Doshi · Machinecraft · 10 min
Nhà máy 100 người không đội data science build company brain Ira: agent chuyên biệt, memory theo tầng, dream cycle ban đêm, không train model.
A Song of Types and Agents
Roberto Stagi · Ratel · 14 min
Vì sao TypeScript vượt Python trên GitHub và đang chiếm application layer của AI agent: coding agents, npm, một codebase, Zod end-to-end.
Semantic Blindness: 500,000 Sensors Confused an LLM
Raahul Singh & Vanč Levstik · Phaidra · 16 min
Vì sao LLM gãy khi phải đọc 500 nghìn tên thiết bị, và cách Phaidra để LLM chỉ lập plan còn code tra cây: 100% chính xác, ít token hơn 300 lần.
The AI bugpocalypse is here. Now what?
Jack Cable · Corridor · 20 min
Frontier model tìm lỗ hổng ngày càng giỏi, nhưng lỗ hổng vẫn thuộc các lớp cũ: loại bỏ cả lớp bằng memory safety, đặt guardrails cho AI coding.
What Does Done Even Mean? Agents and Paperclip's Liveness Model
Dotta · Paperclip · 7 min
Coi done là một object chứ không phải checkbox: tách bó claim, cân liveness với assurance, và các cơ chế control plane của Paperclip cho agent.
ReviewDebt: a practical framework for scoring every pull request
Sachin Gupta · eBay · 25 min
ReviewDebt chấm mỗi PR bằng năm tín hiệu deterministic để đo khoảng trống giữa code coding agent tạo ra và code con người thật sự review.
The UX of AI: Making AI-Powered Apps Your Users Don't Hate
Kathryn Grayson Nanz · Progress Software · 36 min
Năm trụ cột UX cho tính năng AI: trust, clarity, control, transparency, meaningful benefit, kèm pattern thật như citation, action plan, nút stop, undo.
Your Agents Need a Save Button
Hamza Tahir · ZenML · 17 min
Checkpoint state của agent trong một durable runtime để replay run production với model hay tool khác, diff kết quả và quyết định trên cả cohort.
Agent Output Is Not UX: Rendering Layer Your LLM Pipeline Is Missing
Bala Ramdoss · Amazon · 14 min
Lớp generative UI giữa model và màn hình: rendering contract có version, streaming vào typed component, và BFF giữ cho mobile client an toàn.
Can Oncology Workflows Run Without Human Touch?
Anant Shankhdhar · RISA Labs · 17 min
Bốn agent tự động hoá prior authorization cho thuốc ung thư: deterministic check trước, bằng chứng nhiều nguồn để có confidence, reasoning layer cho ca khó.
Build the AI GTM Agent That Knows the Buyer Before the First Message
Dr. Sajjan Kanukolanu · Position² · 26 min
Kiến trúc ba lớp signals, buyer intelligence, action và context graph để AI GTM agent biết người mua trước tin nhắn đầu, kèm bốn chỗ hệ thống gãy.
Don't Let the LLM Drive
Ornella Bahidika & Joel Allou · Microsoft · 7 min
Voice tutor Ace giữ control flow trong harness: bài học là state machine, model chỉ nhận contract hẹp, nhờ đó chạy reliable trên Haiku 4.5.
Enterprise Agents Have a Structure Problem
Ishita Daga · Tesla · 12 min
Data agent doanh nghiệp sai vì thiếu structure chứ không thiếu context: thứ bậc source of truth, context lifecycle có eval, và bài toán preference còn mở.
Medic for Apache Spark: First Aid for Failing Jobs
Drasko Profirovic · Pinterest · 11 min
Pinterest xây agent chẩn đoán Spark job fail: MCP, E2E test harness record/playback, lọc exception, metrics thành hình, rồi multi-agent trên deepagents.
Skills are the New SDKs
Elvin Aghammadzada · DataRobot · 27 min
Vì sao skill là lớp experience mới cho agent: context rot, progressive disclosure, skill vs MCP, cấu trúc SKILL.md và rủi ro của hệ sinh thái skill.
Designing Voice Agents for Real Conversations
Chintan Agrawal & Daniel Wirjo · AWS · 33 min
Turn-taking cho voice agent qua ba level: Silero VAD, turn detection trong STT, rồi VAD cộng Smart Turn; kèm latency budget, LLM TTFT và demo Pipecat.
When Agents Meet Physical Data: The Other Physics of Agent Harnesses
Dmitry Petrov · DataChain · 28 min
Vì sao coding agent hỏng với video và sensor data, và cách dựng data harness bốn phần: see, run, verify, remember, với demo DataChain và Claude Code.
Why Your Agent Disagrees With Itself (And What To Do About It)
Diane Lin · Datadog · 26 min
Agent flip-flop là dấu hiệu của gray zone: dùng disagreement để chọn ca cho người review, rồi thêm semantic và episodic memory thay vì fine-tune.
Your Voice Agent Doesn't Need a Frontier Model
Ornella Bahidika & Joel Allou · Microsoft · 6 min
Voice tutor Ace chọn model nhỏ vì latency budget khoảng 950 ms: state machine và mastery tracking nằm trong code, model chỉ nói, nhờ đó Haiku 4.5 thay được Opus 4.7.