让AI写工厂代码,先跑起来才算数
写代码的AI已经能生成工厂里的PLC程序,但以前只检查代码本身,没人真把它放进工厂系统里跑一遍。这篇把验证门槛抬高:代码必须通过编译、行为检查,还要在真实运行环境里跑出和参考程序一致的轨迹,才算完成。结果在7个模型上,严格通过率平均72.6%,比基线高出一截;最关键的差距出现在动态运行层——静态评分大家只差10分,一跑起来,基线掉到22.4~31.4,它拿到52.2。它不是你明天就能用的工具,但指向一个趋势:AI写工业代码,光看代码已经不够,得看它跑起来的样子。
📄 原文摘要(英文)
Programmable logic controllers (PLCs) run industrial plants, and large language models can already generate independent program organization units (POUs) for them. Whether such logic integrates into an existing PLC project and then runs correctly has been checked only in limited tests. We present SemaPLC, a project-grounded and verification-gated agent harness assembled from conventional tools but governed by a strict completion rule. Rather than stopping when the model judges its own output adequate, SemaPLC declares a task complete only when logged external checks confirm it. Those checks cover the specification, the compilation, and the behavior on a live runtime. On 117 independent-POU tasks matching existing benchmarks, it attains the highest strict verified pass rate on all seven models (72.6\% mean). On a project-context track of 65 tasks whose generated logic must compile and run inside a real project, it attains the highest mean on integrated compilation, static behavior, and dynamic behavior. Of the three layers, dynamic behavior is the most revealing. We measure it by deploying the generated and the reference logic to a live PLC runtime and comparing their executed traces. All methods fall within 10 static points of one another, whereas dynamic scores separate them sharply, from 22.4 to 31.4 for the baselines against 52.2 for SemaPLC. Overall, our verification-gated harness raises the mean at every layer and most sharply at runtime. Execution, not static scoring, is the faithful test of whether generated control logic actually works. SemaPLC is open-sourced at https://github.com/midea-ai/SemaPLC.