把AI代理写成Python对象,NVIDIA新框架让调试像改代码一样简单
传统AI代理开发要写一堆模板、工具描述、回调函数和工作流图,像拼乐高但零件散落一地。NVIDIA的新框架NOOA把整个代理变成一个Python对象:它的方法就是AI能调用的动作,字段是状态,文档字符串是提示词,类型注解是契约。最妙的是,方法体写'...'的部分由AI在运行时自动填充,有正常代码体的部分则保持确定性执行——开发者和AI用同一套接口,测试、追踪、重构都跟普通软件一样。它不是你明天就能用上的,但如果你被代理的不可控和调试痛苦折磨过,这个思路值得关注。
📄 原文摘要(英文)
Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA takes a simpler approach: an agent is a Python object. Its methods are the actions the model can take, fields are its state, docstrings are its prompts, and its type annotations are contracts. A method whose code body consists of "..." is completed at runtime by an LLM-driven agent loop, while methods with normal bodies remain standard deterministic Python. This gives developers and agents the same interface, so agent behavior can be tested, traced, refactored, and improved just like other software. This paper makes three contributions. (1) We present the agent-as-a-Python-object programming model and the design principles behind it. Where Python has existing abstractions, we adopt them directly. Agent-specific capabilities--context, events, state rendering, long-term memory, and validated LLM loops--are exposed through simple Pythonic APIs, so both developers and agents share one familiar programming model. (2) We identify six model-facing ideas that NOOA is, to our knowledge, the first to combine on a single surface: typed input/output, pass-by-reference over live objects, code as action, programmable loop engineering, explicit object state, and model-callable harness APIs for context and events. We find the community already converging on several of these ideas--often as experimental or partial features--and present the comparison to encourage further adoption. (3) We demonstrate that current models use this interface effectively, both in targeted capability tests and on agentic and reasoning benchmarks such as SWE-bench Verified and Terminal-Bench 2.0 and ARC-AGI-3.