Training terminal agents requires scalable executable supervision, but synthesizing high-quality terminal tasks remains a challenge. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated from inconsistent assumptions, the task may be unsolvable or incorrectly evaluated.
Multi-stage synthesis often discards the goals, dependencies, state transitions, and procedural constraints encoded in the original sources. To address this, researchers at Hugging Face present FACET (Fine-grained Agentic Construction of Executable Tasks), a framework that focuses on both information preservation and cross-artifact consistency.
FACET reconstructs related agent skills into coherent, information-rich scenarios, then realizes and repairs the execution environment before generating the final task artifacts. The resulting container state serves as shared grounding for the instruction, solution, and verifier, while execution-based validation and targeted repair correct artifact-specific failures without unnecessarily regenerating valid components.
The framework produces complex terminal tasks with dense executable checks, and successful trajectories collected from these tasks provide effective, data-efficient supervision. Fine-tuning models across multiple scales consistently improves performance on Terminal-Bench 2.1, and analyses of alternative generation schemes support the importance of environment-grounded construction for task validity and solution-verifier alignment.
These results establish source-intent preservation and shared executable-state grounding as key principles for scalable terminal-task synthesis.