Large language models (LLMs) are increasingly used to automate data-processing workflows, but coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. This disconnect, termed the NL2Pipeline gap, is addressed by a new platform called DataFlow-Harness, introduced by researchers at Hugging Face.
DataFlow-Harness guides an LLM agent to construct platform-native directed acyclic graphs (DAGs) through typed, incremental mutations rather than free-form scripts. The platform combines three key components: DataFlow-Skills for procedural guidance, a Model Context Protocol (MCP) layer that exposes the live operator registry and current pipeline state, and DataFlow-WebUI, which synchronizes conversational authoring with a visual DAG editor.
On a 12-task data-engineering benchmark, DataFlow-Harness achieves a 93.3% observed end-to-end pass rate. Relative to Vanilla Claude Code, it reduces measured monetary cost by 72.5% and generation latency by 49.9%. Its observed pass rate is within 0.9 percentage points of the Context-Aware Claude Code baseline while its cost is 42.8% lower.
Per-task analysis indicates that Skills are most useful when construction depends on implicit procedural knowledge. These results show that live platform grounding can produce persistent, editable workflow artifacts with reliability close to script-generation baselines and with lower construction cost and latency.