SFML

Software Factory Markup Language (SFML)

Version: v0.1 - Status: Draft - License: MIT


Introduction

SFML is an open standard for defining how agents and humans interact. Sometimes we call this a software factory. At its core, SFML is a graph of steps where each step is a human or agentic process. SFML graphs support loops and deeply constrain branching to ensure input state to any step is perceivable via a simple reading of the graph.

The motivation behind SFML is to separate orchestration of process from agent runtime while also providing organizations a reviewable artifact so that individual, team and company processes and be reviewed and improved.


1. Scope

1.1 What SFML defines

1.2 What SFML does not define

2. Normative references

3. Terms and definitions

3.1 Notational conventions

The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 (RFC 2119, RFC 8174) when, and only when, they appear in all capitals as shown here.

3.2 Terms

3.2.1 Factory

A document, conforming to this specification, that declares a graph of steps and the routing between them. A factory is a static artifact; it describes work, not any particular execution of it.

3.2.2 Run

One execution of a factory, identified by a run id minted by the implementation. A run accumulates FactoryState (§9.2) as its steps produce results.

3.2.3 Branch

A live line of control within a run. A run has exactly one branch except while control is inside a parallel step, where it has one branch per child (§9.8). A branch is at all times running, awaiting_input, errored, or done (§11.1).

3.2.5 Iteration

A single traversal of a step from entry to success, counted toward a step's max_iterations. Where this document describes a step entry whose outcome is not yet known — the traversal is still in progress, or it ended in an exception rather than success — it says "entry" or "entering," not "iteration": only a successful entry is an iteration, per §9.6.

3.2.6 Attempt

A single try at executing a step within one entry. An attempt that fails does not append to results; an entry may consist of several attempts, ending in one iteration, when retry (§6.10) applies.

3.2.7 Prompt template

The content named by an agent step's prompt_path or prompt (§6.5): text interspersed with placeholders (§7.9), rendered against that step's PromptVars (§7.5.2, §9.9) to produce the text sent to the harness.

3.3 Roles

3.3.1 Author

The person or persons, or the system, that writes a factory file.

3.3.2 Caller

The person or system that starts a run, supplying the values bound to parameters (§6.3), or that resumes a blocked run by supplying a human step's result, granting additional iterations, budget, or by supplying a caller-authored result in place of a wedged agent step (§11.4). This specification does not distinguish whoever starts a run from whoever later resumes one; both act through the same addressed calls (§9.1, §11.3), and a conforming implementation MAY apply its own access control to either without SFML's involvement.

3.3.4 Harness

The component that executes an agent step and returns a result. This specification names the harness, bounds what it must report to a conforming implementation (cost, retryability of a failure, and continuity of a session across a resume), and does not describe how it works internally.

3.3.5 Implementation

Software that parses, lints, or runs factory files in conformance with this specification.

4. Conformance

4.1 Conformance classes

4.1.1 Parser

A Parser converts a document into the data model of clause 6, or rejects it. Its question is whether every part of the document maps to something this specification defines, the way a reader of a binary format rejects a byte sequence that maps to no known structure. A Parser rejects a document whose bytes are not valid UTF-8, whose syntax is not valid YAML, that has a duplicate key, that has a field this specification does not define at that position, that lacks a required field, or that has a value not of its field's declared type (§4.3). A Parser treats an expression and a prompt template as opaque strings. It does not resolve one name in the document against another, and it reads only the document itself.

4.1.2 Linter

A Linter accepts a factory a Parser has produced and checks the rules that depend on relationships between its parts or on the content of its strings and files: that names resolve (§8.2), the graph's routing, reachability, and termination (§8.3–§8.5), expressions (§7.1, §7.8, §8.6), and prompt templates, including the files prompt_path names (§7.9). It reports the stable identifier (§8.7) of each violated rule. A Linter never sees a document its Parser rejected, and does not execute a run.

4.1.3 Runner

A Runner executes runs of a factory that has passed linting. It owns what an agent step must achieve regardless of which harness it binds to: reporting cost (§9.7), continuing the same session across a resume (§11.7), and classifying a failure as retryable or not (§10.3). How it obtains any of these from a particular harness is the Runner's own business, not something this specification constrains.

4.2 Requirements by class

Requirement Parser Linter Runner
Reject every document that does not map to the clause 6 data model (§4.3) MUST MUST MUST
Produce the clause 6 data model from a document that does MUST MUST MUST
Report every Linter rule's violation with its §8.7 identifier (§4.3) — MUST MUST
Refuse to start a run of a factory that fails linting (§4.3) — — MUST
Implement admission (§9.1) — — MUST
Implement the execution model of clause 9 — — MUST
Raise the exception classes of clause 10 under their stated conditions — — MUST
Implement resume addressing and payloads of clause 11 — — MUST
Satisfy the resumability requirement of clause 12 — — MUST

A single piece of software MAY implement more than one conformance class. An implementation that claims the Runner class MUST also satisfy the Parser and Linter requirements, since a Runner MUST refuse to start a run of a factory that fails linting.

4.3 Division of checks between Parser and Linter

Every static rule in this specification belongs to exactly one of the two classes. The test is what the rule needs to see:

Checked by the Parser Checked by the Linter
UTF-8 encoding and YAML syntax (§5.1) start and every to name a declared step (§8.2)
No duplicate keys (§5.6) Routing is total (§8.3)
No field this specification doesn't define at that position (§5.4) Every step is reachable and reaches a result (§8.4)
Every required field present (§6.2–§6.9) Every cycle is bounded (§8.5)
Every value of its field's declared type (§6.1) Every expression is valid CEL, within §7.2's grammar (§7.1, §7.8)
No field on a step type it doesn't apply to (§6.4) Every reference resolves and is reachable, in its binding environment (§8.6)
Exactly one of prompt and prompt_path (§6.5) The file prompt_path names can be read as UTF-8 (§7.9)
A parallel child is an agent or human step with no next (§6.7) Every prompt template's placeholders are well-formed (§7.9)

A Parser reports no identifiers: this specification defines only whether a document is accepted. A document the Parser rejects is never linted, so a Linter rule never needs to handle a document that doesn't map to the data model.

4.4 Precedence of this document over Annex A

Annex A provides a JSON Schema for the data model of a factory document, covering most Parser rules (§4.3). Where the schema and the normative text of this document disagree, this document governs. The schema is provided to make the Parser's checks cheap to implement; it is not a complete Parser (it cannot detect duplicate keys, and leaves the two-decimal-place limit on Decimal USD to the Parser, since common validators test it with binary floating point), and it is not a substitute for the Linter rules, which need reasoning across the document that a schema validator cannot perform.

5. Document format

5.1 Encoding and surface syntax

A factory document is a YAML 1.2 document, encoded as UTF-8. Every value in the data model of clause 6 that is legal YAML MUST be representable in a factory document; an implementation MAY also accept the equivalent JSON document, since every JSON document is valid YAML.

Comments carry no meaning: an implementation MUST NOT assign meaning to a comment, and a document's data model is the same with or without them. This specification does not require a tool that rewrites a factory document to preserve its comments.

A factory document stored as a file SHOULD use the .sfml extension (for example, checkout.sfml), optionally combined with the underlying syntax as .sfml.yaml or .sfml.json. This is a naming recommendation, not a validity requirement: a document named with a bare .yaml, .yml, or .json extension remains a conforming factory document, and an implementation MUST NOT reject a document solely because of its file extension.

5.2 Names and identifiers

A StepName is a non-empty string. It MUST NOT contain ., which is reserved as the qualifier separator between a parallel step and its child (§5.3). StepName comparison is case-sensitive: Plan and plan name distinct steps. A Name used as a parameters or prompt_vars key is subject to the same case-sensitivity rule.

This specification does not further restrict the character set of a StepName or a Name beyond excluding .; an implementation MAY impose additional restrictions of its own (for example, to fit an identifier into a storage system) provided it documents them.

5.3 Qualified names

A step declared as a child of a parallel step (§6.7) is addressed, outside its own step definition, by its qualified name: <parallel>.<child>, where <parallel> is the name of the enclosing parallel step and <child> is the child's own name. Qualified names are what a resume (§11.3) addresses and what an exception raised inside a region names. A child's name need only be unique within its own parallel step; two different parallel steps may each declare a child named lint without conflict, since their qualified names differ.

5.4 Unknown fields

An implementation MUST reject a document containing a field not defined by this specification, at every level of the data model, including the factory's top level, every step, every connection, and every nested object this specification defines. This is a hard error, not a warning, and it applies regardless of whether the unrecognized field's name resembles a future or vendor-specific extension.

5.6 Duplicate keys

A YAML or JSON mapping with a duplicate key at any level of a factory document (for example, two steps entries with the same StepName, or a step object with two budget fields) is a parse error. A conforming Parser MUST reject such a document rather than silently applying "last value wins" or any other resolution.

5.7 Graph size

This specification sets no maximum on the size of a factory. A conforming implementation MUST NOT reject a factory for the number of steps, connections, or parallel children it declares.

6. Data model

6.1 Value types

Type Description
String A YAML/JSON string.
Integer A YAML/JSON integer with no fractional component.
Decimal USD A non-negative decimal number of United States dollars with at most two decimal places (§9.7).
JSON Schema A value conforming to the JSON Schema specification referenced in clause 2.
Expression A String holding an SFML expression (clause 7). A Parser checks only that it is a String; whether its content is a valid expression is a Linter rule (§4.3, §7.1). Two flavors exist: FactoryState Expression and PromptVars Expression (§7.5), distinguished by the binding environment they resolve against, not by syntax.
StepName A name identifying a step, per §5.2.
SFMLVersionString A string of the form v<major>.<minor>, e.g. "v0.1".

6.2 Factory

The top-level object of a factory document.

Field Type Required Notes
sfml SFMLVersionString yes The version of this specification the document targets.
description String no Free text an implementation MAY display to its users.
parameters Record<Name, JSON Schema> no The factory's signature (§6.3).
start StepName yes The entry point. It is never inferred from graph topology, since a factory may loop back to its first step.
steps Record<StepName, Step> yes The factory's steps (§6.4–§6.8).
assignee String no The run's default DRI. Opaque to SFML (§6.11).
budget Decimal USD no The ceiling across every agent step of every iteration in the run (§9.7).

Unknown fields MUST be rejected (§5.4).

6.3 Parameters

parameters is the factory's signature: the contract between a factory and whatever starts a run of it. Each entry is a JSON Schema.

6.4 Step: common fields

Every step has a type of agent, human, parallel, or result. The following fields are common to two or more of these types; a field used by exactly one step type is documented in that type's own subclause (§6.5–§6.8) instead.

Field Type Required Applies to
type enum agent | human | parallel | result yes all
description String no all
next List<Connection> yes agent, human, parallel
max_iterations Integer no agent, human, parallel
result_schema JSON Schema yes agent, human

A field present on a step of a type it does not apply to MUST be rejected (§5.4). This applies both to the fields above and to the type-specific fields documented in §6.5–§6.8.

result_schema, where present, validates the object an agent or human step produces, before routing (§9.5) is evaluated for that step. A parallel step MUST NOT declare result_schema: its result is computed from its children, not authored (§6.7).

6.5 Agent step

An agent step invokes a harness loop once per attempt (§3.2.6) and produces a result validated against result_schema.

Field Type Required
harness <name>[@<version>] yes
harness_config Record<Key, Any> no
prompt_path | prompt path | String yes (one of)
prompt_vars Record<Name, Expression> no
budget Decimal USD no
retry Integer no

6.6 Human step

A human step blocks its branch (§11.1) until a caller supplies a payload validated against result_schema.

Field Type Required
assignee String no
instructions Expression no

6.7 Parallel step

A parallel step declares a map of named child steps in its own steps field, runs them concurrently, and joins once every child has produced a result.

Field Type Required
steps Record<Name, Step> yes

6.8 Result step

A result step is terminal: reaching one ends the run, and a run that has reached one MUST NOT be resumed or restarted.

Field Type Required
outcome enum complete | terminal_failure yes
value Expression no

A result step MUST NOT declare next, result_schema, harness, or any field specific to another step type.

value, where present, is a FactoryState Expression (§7.5.1), evaluated against FactoryState (§9.2) to produce the step's result. Where value is omitted, the step's result defaults by outcome:

6.9 Connection

A Connection is one entry of a step's next list.

Field Type Required Notes
when Expression no Absent means unconditional. There is no else key; the fallback is a connection with no when.
to StepName yes Exactly one target. Fan-out is expressed only by a parallel step (§6.7), never by a connection.

6.10 Retry

retry, on an agent step, is an Integer: the maximum number of attempts within one entry to the step. Where present, it MUST be at least 1, since an entry always makes at least one attempt; a document declaring a smaller retry MUST be rejected. It governs re-attempting a harness failure that the harness itself has classified as retryable (§10.3); an implementation MUST NOT re-attempt a failure the harness has not classified as retryable, regardless of retry. An absent retry means one attempt: a retryable failure on that single attempt raises harness_error (§10.3) immediately.

Spacing between attempts is not an authored field: where a harness states a backoff (§10.3), an implementation SHOULD honor it as best practice.

A failed attempt does not append to results (§9.3); exhausting retry's attempts, or encountering a non-retryable failure, raises harness_error (§10.3). Every attempt within one entry continues the same harness session (§11.7).

A resume from harness_error (§10.3) or schema_violation (§10.4) that re-runs the step opens a new entry: retry's attempt count applies fresh to that entry, exactly as it did to the entry the exception was raised from. That new entry still continues the resumed entry's harness session (§11.7).

6.11 Assignee

assignee is declared at two levels: on the Factory (§6.2) and on a human step (§6.6). Both exist so an Author can record, in the file, who this specification calls the DRI of the run or the worker on a step — in whatever terms the implementation's own systems use for ownership.

The value of either field is a String, opaque to SFML. This specification does not define what it denotes — a person, a team, a rotation, a queue — and does not define how, or whether, an implementation acts on it. assignee is a place for an Author to record ownership; it is not a mechanism SFML uses to resolve, notify, or route to anyone.

6.12 Harness reference and harness configuration

harness, REQUIRED on an agent step, is a String of the form <name>[@<version>] naming the harness that executes the step. Resolution of <name> and <version> to an executable harness is implementation-defined, except that resolution MUST be deterministic for a given implementation and configuration. Where a step's harness fails to resolve, a run using that step MUST NOT be started; an implementation MUST reject it at admission (§9.1).

harness_config is an OPTIONAL record of harness-defined keys and values, passed through to the named harness unexamined by the rest of this specification. Everything about a run's workspace — repository, branch, worktree, and what an agent may read or write — is confined to harness_config; portability of a factory file, as promised by this specification, ends at this boundary.

7. Expression language

7.1 Relationship to CEL

Every SFML expression MUST parse as valid CEL. An implementation MAY use a CEL evaluator with the function library of §7.4 registered as extensions, or MAY implement the subset of CEL this clause defines directly in a language with no usable CEL binding; either satisfies this clause provided it accepts exactly the grammar of §7.2 and rejects everything else.

Checking an expression is a Linter rule, at every expression site and in every prompt template placeholder (§7.9). A Linter MUST reject an expression that does not parse as CEL, reporting invalid-expression (§8.7). An expression that parses as CEL but uses a construct outside §7.2 is instead reported as prohibited-expression-construct (§7.8), so a Linter must recognize CEL's syntax, not only the subset, to tell the two apart.

7.2 Grammar subset

An SFML expression is drawn from a closed, small subset of CEL: field selection, indexing, comparison and boolean operators, literals, and calls into the standard function library of §7.4. User-defined functions, arithmetic on step results, and collection macros are not part of this grammar; an expression using any of them is not a valid SFML expression regardless of its validity as general CEL.

7.3 Null semantics and propagation

There is no optional-chaining operator. Field access on null yields null rather than an error; combined with last() (§7.4) yielding null over an empty list, an expression such as last(emptyList).some_field is valid and evaluates to null.

Selecting a field that a non-null object does not have is not null: it is a runtime evaluation error (§7.7). An expression should only read fields that the result_schema or parameter schema it resolves against guarantees are present, so a missing field means a defect in the factory or in the data it was given, not a value to route on. Where an Author needs to branch on an optional field, the schema should make it required and nullable instead.

7.4 Standard function library

The standard function library is closed. The following CEL extensions are REQUIRED of every conforming implementation, and no others allowed:

Function Behavior
last(list) The last element of list, or null if list is empty.
empty(x) true if x is an empty list, empty string, or null; false otherwise.
notEmpty(x) The negation of empty(x).

7.5 Binding environments

An SFML expression is evaluated against exactly one binding environment, determined by where it appears in the data model. Reaching outside the environment bound to a given site is a static error (§8.6).

7.5.1 FactoryState expressions

A FactoryState Expression targets FactoryState (§9.2). It is the kind of expression used in a Connection's when (§6.9), a human step's instructions (§6.6), a result step's value (§6.8), and an agent step's prompt_vars (§6.5).

last_result (§9.2) is reachable at all four of these sites: it is an ordinary part of FactoryState, and nothing about it changes which sites are FactoryState Expression sites.

7.5.2 PromptVars expressions

A PromptVars Expression targets the PromptVars (§9.9) of the agent step whose prompt template it appears in. FactoryState is not reachable from this environment; anything a template needs MUST first be named in that step's own prompt_vars.

7.6 Type checking

An implementation SHOULD statically type-check an expression against the schema of the values it resolves against — result_schema of the steps it references, and the JSON Schema of any referenced parameter — and SHOULD report a type error before a run starts rather than at evaluation time.

7.7 Evaluation errors

An expression that is well-typed per §7.6 but fails at evaluation time (for example, indexing past the end of a list), or a prompt template (§3.2.7) that fails to render against its evaluated PromptVars, is a runtime evaluation error. An implementation MUST raise an exception rather than silently producing a value or a partial rendering: routing_error (§10.8) where the failure is evaluating a Connection's when, or expression_error (§10.7) for every other evaluation site. These are separate classes, not one, because the two failures are resumed past differently (§11.4).

7.8 Prohibited constructs

An expression MUST NOT contain a user-defined function, arithmetic on a step result, a collection macro, or any construct outside the grammar of §7.2, even where the underlying CEL implementation would otherwise accept it. A Linter MUST reject such an expression, reporting prohibited-expression-construct (§8.7). Like the checks of §8.6, this needs only the expression's parse tree, never its evaluation.

7.9 Prompt template placeholders

A prompt template — the content of prompt, or of the file prompt_path names — MUST be encoded as UTF-8, independent of the factory document's own encoding (§5.1), so that the delimiter below is unambiguous.

A placeholder in a prompt template (§3.2.7) is delimited by double guillemets at the start and end of the expression. For example: «« expression »». A single guillemet is ordinary literal text and MUST NOT be treated as part of a placeholder; only the doubled pair opens or closes one.

This specification defines no escape for a literal «« or »» in template text. A single « or » is unaffected by this rule and needs no special treatment; an author who needs the doubled sequence itself to appear literally, rather than open a placeholder, has no way to express that.

Checking a prompt template is a Linter rule (§4.3). For every agent step, a Linter MUST:

8. Graph validity

8.1 Validation stages

Validity is checked in three stages, each with a different reach: parse (§4.1.1, whether the document maps to the data model of clause 6), lint (§4.1.2, the rules of this clause and of §7.1, §7.8, and §7.9 across a parsed factory and the prompt files it names), and admission (§9.1, checks that require information outside the factory, such as an identity system). §4.3 assigns every static rule to exactly one of parse and lint. Each stage runs only on what the one before it accepted: a document that fails to parse is never linted, and a factory that fails lint is never admitted. A document that passes lint is not thereby known to be admittable, and the stages MUST be kept distinct by a conforming implementation.

8.2 Step references

The data model gives start and a Connection's to the type StepName (§6.2, §6.9), which a Parser checks. Whether the name is declared is a relationship between two parts of the document, so it is a Linter rule. A Linter MUST enforce that every step referenced by start or by a Connection's to is declared in steps.

8.3 Totality of routing

This is a Linter rule, not a schema rule: it constrains the last item of an ordered list positionally, which JSON Schema (Annex A) can only express awkwardly and which is, in any case, the class of graph-level reasoning §4.4 reserves for clause 8 rather than for the schema. Its name describes the property it establishes — that routing is total — not the mechanism.

The last Connection of every non-result step MUST omit when (§6.9). A Linter MUST reject a step whose next list does not end this way. This is what makes the reachability check of §8.4 sound: since every step always has somewhere to go, "no path forward" can only be a lint-time defect, never a runtime condition to detect.

8.4 Reachability

Every step MUST be reachable from start, and every path from start MUST reach a result step; a Linter MUST reject a factory violating either condition.

A Linter MAY check this syntactically: ignore when semantics, treat every Connection as traversable, and rely on totality of routing (§8.3) to make graph reachability the correct check. This is sufficient and is the minimum a conforming Linter must do. A Linter MAY instead perform a semantic analysis of when conditions to detect a step that is reachable syntactically but unreachable given what its guards can evaluate to; doing so MUST NOT cause it to accept a factory the syntactic check would reject.

8.5 Termination

Every cycle in the graph MUST be bounded: a Linter MUST reject a cycle containing no step with a finite max_iterations, since such a cycle has no syntactic guarantee of ever reaching a result step.

8.6 Reference and binding validity

Every check in this subclause is a walk of an expression's parse tree looking for field-selection chains (results.X, parameters.<name>, and so on) — the kind of structural traversal any CEL AST exposes without evaluating the expression. None of them require evaluating the expression itself, and none require more of a CEL implementation than parsing to an AST already does.

This subclause governs references and binding environments only. A field's own value type — for example, that budget (§9.7) has at most two decimal places, or that assignee (§6.11) is a String — is validated as part of the data model of clause 6 rather than checked here.

last_result (§9.2) is a FactoryState identifier like parameters and results, but since it does not name a step or a parameter, none of this subclause's checks apply to a reference to it.

8.7 Diagnostics and error identifiers

Every Linter rule (§4.3) — the rules of this clause and of §7.1, §7.8, and §7.9 — has a stable, unique identifier that a conforming Linter MUST report on failure. Parser rules have none. Message text accompanying an identifier is implementation-defined. This subclause holds the registry of identifiers for this specification; an identifier, once assigned, MUST NOT be reused or renumbered by a later version of this document.

Identifier Clause Rule
unknown-step-reference §8.2 start or a Connection's to names a step not declared in steps.
non-total-routing §8.3 The last Connection of a non-result step does not omit when.
unreachable-step §8.4 A step is not reachable from start.
no-path-to-result §8.4 A path from start does not reach a result step.
unbounded-cycle §8.5 A cycle contains no step with a finite max_iterations.
unknown-parameter-reference §8.6 An expression references parameters.<name> for an undeclared <name>.
unknown-step-result-reference §8.6 An expression references results.<name> for an undeclared <name>.
unreachable-reference §8.6 An expression references results.X from step Y with no path X → … → Y.
binding-environment-violation §8.6 An expression references a value outside the binding environment bound to its site (§7.5).
invalid-expression §7.1 An expression, or a prompt template placeholder's enclosed text, does not parse as CEL.
invalid-prompt-template §7.9 A prompt template has a «« that no »» closes.
prompt-file-unreadable §7.9 No readable file is at the path an agent step's prompt_path names.
prompt-file-not-utf8 §7.9 The file an agent step's prompt_path names is not valid UTF-8.
prohibited-expression-construct §7.8 An expression uses a user-defined function, arithmetic, a collection macro, a function outside §7.4, or any other construct outside the grammar of §7.2.

8.8 What lint cannot check

Lint operates on the graph's shape and cannot evaluate expression semantics, cannot know whatever an implementation checks assignee (§6.11) against, and cannot know what a harness will report about cost or retryability at runtime. A rule that would require any of these is not a lint rule; where such a rule exists in this specification, it is checked at admission (clause 9) or raised as a runtime exception (clause 10) instead.

9. Execution model

9.1 Admission

Before a run is given an id, an implementation MUST:

An implementation MAY perform additional checks of its own at admission — for example, against assignee (§6.11) or against its own identity system — but this specification imposes none beyond the two above.

A run rejected at admission has neither a run id nor recorded results, and raises no exception class of clause 10; rejection at admission is distinct from, and precedes, everything clause 10 describes.

9.2 FactoryState

FactoryState is the value a FactoryState Expression (§7.5.1) resolves against.

Field Type Notes
parameters Record<Name, Any> The run's parameter values, validated and defaulted at admission. Constant for the run's lifetime.
results Record<StepName, List<StepResult>> Oldest to newest per step. last() (§7.4) retrieves the most recent.
last_result Any The result of whichever step's fired Connection (§6.9) most recently routed into the step whose expression is being evaluated. null at start (§9.1), which has no incoming Connection. Each child of a parallel step (§6.7) sees the same last_result the parallel step itself saw, not the parallel's own (not yet produced) result; the step a parallel routes to afterward sees the parallel step's combined StepResult as last_result, per the general rule above. Contextual to the expression site's own step (§7.5.1), not a single run-global value. Not subject to §7.6 static type-checking: its static type is Any.

9.3 StepResult

A StepResult is the validated result object a step produces. last(results.review).approved reads approved directly off the object the review step produced.

A StepResult is appended to results only on the success of an entry to a step (§3.2.5); a failed attempt appends nothing. Each step type defines how its own result is created: see §6.5 for an agent step, §6.6 for a human step, §6.7 for a parallel step, and §6.8 for a result step.

9.4 Step lifecycle

On each entry to a step, an implementation:

  1. Checks the step's iteration bound (§9.6); if exceeded, raises iteration_limit (§10.5) without starting the step.
  2. For an agent step, evaluates prompt_vars against FactoryState, renders the prompt as PromptVars (§9.9), and invokes the harness for one or more attempts per retry (§6.10), tracking reported cost against the step's budget (§9.7).
  3. For a human step, blocks the branch in state awaiting_input (§11.1) until a valid resume arrives.
  4. For a parallel step, starts every child concurrently (§9.8) and blocks until all have produced a result.
  5. Validates the produced payload against result_schema, where the step type declares one; a validation failure on an agent step raises schema_violation (§10.4).
  6. Appends the resulting StepResult to results and evaluates routing (§9.5).

9.5 Routing

On an iteration of a non-result step, an implementation evaluates the step's next list in declared order (§6.9) and transitions control to the target of the first Connection whose when is absent or evaluates to true. Routing always selects exactly one target.

9.6 Iteration bounds

max_iterations is a step field, checked on arrival:

On entering step X, if X has already run max_iterations times, the step does not start.

9.7 Budgets and monetary arithmetic

budget is a Decimal USD value, at most two decimal places (§6.1). 5.00, 0.25, and 12 are legal values; 1.005 is not a legal Decimal USD value and MUST be rejected wherever budget appears.

9.8 Concurrency and join

Concurrency takes exactly one shape, the parallel step (§6.7). On entering a parallel step, an implementation starts every declared child concurrently, opening one branch per child (§11.1). The step joins — appends its result and resumes single-branch control — once every child has produced a result.

9.9 Prompt rendering

An agent step's prompt_vars (§6.5) is evaluated as a map of FactoryState Expressions before the step's harness is invoked. The resulting bindings form PromptVars:

Field Type Notes
prompt_vars Record<Name, Any> The evaluated prompt_vars of the agent step.

A prompt template therefore references prompt_vars.issue, not a bare issue: the namespace lets a future version of this specification add a sibling field to PromptVars without making an existing template's bare names ambiguous. FactoryState is not reachable from a prompt template; anything a template needs MUST be named first in the step's own prompt_vars (§7.5.2).

Rendering means substituting every placeholder (§7.9) in the template's text with the string produced by evaluating it against PromptVars; the text sent to the harness is the result. A parse or evaluation failure in any placeholder raises expression_error (§10.7) before the harness is invoked for that entry.

9.10 Termination

A run that reaches a result step is terminal. Its outcome and value are recorded per §6.8, and the run MUST NOT be resumed or restarted; every branch that was live is now done (§11.1).

9.11 Determinism boundary

An agent step's invocation of its harness is the sole point of nondeterminism in a run: the harness's output for given inputs is not guaranteed to repeat. Every other part of execution — routing (§9.5), iteration accounting (§9.6), budget accounting (§9.7), and join ordering (§9.8) — is a pure function of the StepResults already in FactoryState and the file's own declarations. Consequently, a routing trace (the ordered sequence of step entries and route decisions) produced from a given sequence of StepResults is fully determined; it does not depend on wall-clock time, on the order in which unrelated events occur, or on anything the harness did once its result was reported.

10. Error model

10.1 Exceptions are not routes

A step failure is a factory exception, never a route. on_error and on_timeout fields do not exist in this specification's data model (§6.9); an exception suspends a branch (§11.1) rather than selecting a Connection.

10.2 Exception classes

Class Raised when
schema_violation An agent step's output fails result_schema validation, or yields no value to validate, and retry, where configured, has not fixed it.
harness_error The harness fails to produce a result for a reason outside the agent's own output contract.
iteration_limit Entering a step would exceed its effective iteration bound (§9.6).
budget_exceeded A step's reported consumption reaches or exceeds its effective budget (§9.7).
expression_error A well-typed expression, or a prompt template, fails at evaluation or render time, anywhere but a Connection's when (§7.7, §10.7).
routing_error A well-typed Connection's when fails at evaluation time, after the step's own result already exists (§7.7, §10.8).

All six classes are resumable (clause 11). Nothing in v0.1 is inherently fatal except reaching a result step (§9.10). expression_error and routing_error are two classes rather than one because a step's own result exists when the second is raised and does not when the first is, which is why they take different resume payloads (§11.4).

10.3 Harness failure classification and retry

harness_error deliberately has no enumerated causes. Rate limits, authentication failures, network faults, capacity limits, and a harness's own execution timeouts are all one class, since this specification cannot enumerate what any of them look like across harnesses, and a fixed cause list would be wrong for the next harness bound to it.

10.4 Schema violation

schema_violation applies only to agent steps, even though result_schema is also declared on human steps. A human step's payload is instead validated when a resume arrives, and an invalid payload is rejected at the call (§11.5), so the resume fails, the step remains awaiting_input, and nothing is appended. There is neither a failed attempt to record nor an exception to resume from in this case, because the run never left the state it was already in.

How a harness turns an agent's output into the value validated against result_schema (a native structured-output feature, parsing a fenced block out of the agent's final text, or any other means) is the harness's own business (§3.3.4). When the harness has produced output but no value can be obtained from it, the attempt MUST be treated as failing result_schema validation: it raises schema_violation, not harness_error, and is subject to the retry and reprompt behavior below. harness_error (§10.3) remains the class for a harness that fails to produce output at all.

10.5 Iteration limit

iteration_limit is checked on arrival (§9.6), so nothing has run when it is raised and the state at that point is unchanged. A resume from iteration_limit carries a payload: a non-negative integer number of additional iterations to grant the step. The step's effective bound becomes effective_limit(X) = X.max_iterations + granted(X).

10.6 Budget exceeded

budget_exceeded is raised when a step's reported consumption reaches or exceeds its effective budget (§9.7). A resume from budget_exceeded carries a payload: a Decimal USD amount, at most two decimal places, to grant. The effective ceiling becomes effective_budget(X) = X.budget + granted_budget(X).

A resume from budget_exceeded uses the same harness session that was running when the ceiling was reached, per §11.7.

The address a resume uses depends on the scope that was exceeded and where control was:

In every case, the resume address says where execution continues, while the exception itself says what was exceeded; these need not be the same scope, as the factory-level cases show.

10.7 Expression error

expression_error is raised when an expression (§7.7) — a FactoryState Expression or a PromptVars Expression — fails at evaluation or render time despite being well-typed, anywhere except evaluating a Connection's when (§10.8): an agent step's prompt_vars (§6.5) or prompt template (§3.2.7), a human step's instructions (§6.6), or a result step's value (§6.8). Every site this class covers is evaluated before the step's own result exists, so nothing has been appended to results for the entry (§9.3, §11.1), exactly as for any other exception raised before a step's own entry completes.

A resume re-attempts the entry from the start. A resume MAY instead carry a result in place of re-attempting: an object validated against result_schema for a step type that declares one (agent, human), or, for a result step, which declares no result_schema (§6.4), any JSON value. Either way the supplied value is appended as the step's StepResult, exactly as the escape hatches of §10.3 and §10.4 work. A caller does not need to know which of prompt_vars, the prompt template, instructions, or value actually failed to know what a resume here accepts: the payload is always a stand-in for the step's whole result, regardless of which expression inside that step's entry raised the exception.

10.8 Routing error

routing_error is raised when a Connection's when (§6.9) fails at evaluation time despite being well-typed. Routing is evaluated once the step's own result has already been appended (§9.4, §9.5), so unlike expression_error, the step's StepResult already exists and is not reconsidered; only the routing decision is missing. This is a separate class from expression_error, not a special case of it, precisely so that a caller can tell from the class alone which payload applies (§11.4).

A resume re-evaluates the when list from the start. A resume MAY instead carry a payload naming one StepName from that step's own declared next list, taken as the routing decision in place of re-evaluating when; an implementation MUST reject a payload naming a target the step's next does not declare (§11.5).

11. Pause and resume

11.1 Branch states

Stopping is a property of a branch, not of a run as a whole. A run has a set of live branches — except while control is inside a parallel step, where it has one per child (§9.8). At any time a branch is running, awaiting_input, errored, or done.

A parallel step with four children, one of them a human step still waiting, is a run with one awaiting_input branch and three done or running branches; there is no separate notion of a partially blocked run.

A factory-level budget_exceeded (§10.6) reached inside a parallel step's region is the one exception to "one branch per child": every child that has not yet produced a result collapses into a single errored branch, addressed by the parallel step's own name rather than by any child's qualified name, until a resume there re-forks them back into their own running branches. A child that already produced a result before the region stopped keeps its own done branch, unaffected.

11.2 Derived run status

A run's status is derived, never stored independently of its branches:

  1. running, if any branch is running.
  2. Otherwise errored, if any branch has that status.
  3. Otherwise awaiting_input, if any branch is waiting on input.
  4. Otherwise terminal, once a result step has been reached (§9.10).

11.3 Resume address

A run has an id, minted by the implementation; its type and encoding are unconstrained by this specification.

A resume is addressed to (run_id, step_name), where step_name is the name declared in the file, qualified as <parallel>.<child> (§5.3) for a step inside a region. One address form serves both a human step's awaiting_input block and an exception's errored block; what the resume does follows from the state of the branch it addresses, not from a mode the caller selects.

A resume addressed to a step whose branch is neither awaiting_input nor errored in that run MUST be rejected.

The address is total: it always names exactly one blocked branch — except a factory-level budget_exceeded inside a parallel step's region (§10.6, §11.1), whose single address, the parallel step's own name, represents every one of that region's still-incomplete children at once. Outside that case, concurrent blocks exist only as children of a parallel step, each child has a distinct name within it (§5.3), and regions do not nest, so no two live branches are ever blocked under the same qualified name regardless of what blocked them.

A conforming implementation MUST document how a run id is obtained and how a blocked run at a named step is resumed; this specification requires no more of the mechanism than that.

11.4 Payloads by branch state

Every class that can be resumed past by re-attempting can also be resumed past by supplying the value the automatic path would otherwise have produced: schema_violation (§10.4), harness_error (§10.3), and expression_error (§10.7) all accept an override of the step's StepResult in place of re-attempting. routing_error (§10.8) is the one exception: since the step's own result already exists by the time it is raised, its override is a routing decision, not a result. iteration_limit and budget_exceeded are different again — they take a grant, not a substitute value, since what is missing is permission to keep spending, not a result (§10.5, §10.6). A caller can tell which shape a given errored branch expects from its exception class alone, without inspecting the run's history for context.

The payload a resume carries depends on what blocked the branch:

Branch state Raised by Payload Effect
awaiting_input a human step object matching result_schema appended as the step's StepResult
errored iteration_limit non-negative integer, additional iterations grant recorded (§11.6); the step is entered
errored budget_exceeded Decimal USD, at most two decimal places grant recorded (§11.6); the step is entered
errored harness_error none, or an object matching result_schema the step re-runs on the same harness session; if a payload is given, that result is appended instead (§11.7)
errored schema_violation none, or an object matching result_schema the step re-runs as a new agent turn on the same harness session; if a payload is given, that result is appended instead (§11.7)
errored expression_error none, or an object matching result_schema (or, for a result step, any JSON value) the entry re-attempts; if a payload is given, that result is appended instead (§10.7)
errored routing_error none, or a StepName from the step's own next the when list re-evaluates; if a payload is given, that target is routed to instead (§10.8)

11.5 Rejected resumes

A payload that does not match the row it is addressed to — an object that fails result_schema validation, a grant of the wrong type, a StepName not among the addressed step's own declared next targets (§10.8), or a resume addressed to a class it does not apply to — MUST be rejected at the call. The branch keeps the state it had, and nothing is appended: this is the same rule §10.4 states for a human step's invalid input, generalized to every row of the table in §11.4. A rejected resume is not an attempt and not a failure; the run has not moved.

A grant payload of zero is legal and is a no-op grant: the step is entered and immediately raises the same exception again, since its effective bound or budget is unchanged. There is deliberately no way to lower a ceiling once raised, only to raise it further: a negative grant, of iterations or of budget, MUST be rejected at the call like any other payload that does not match its row.

11.6 Grants

A grant — additional iterations (§10.5) or additional budget (§10.6) — is not part of FactoryState (§9.2) and is not reachable from any expression. A grant is a record in the run's history (clause 12), consulted only by the effective-limit and effective-budget computations of §9.6 and §9.7.

11.7 Session continuity

A harness session is whatever state a harness holds across invocations, for example the prior turns of a conversation. Every harness invocation either opens a new session or continues an existing one, and a conforming Runner MUST choose as follows:

  1. A new entry opens a new session. The first attempt of an entry into a step (§3.2.5) opens a new session, unless rule 3 applies. This includes an entry made by routing back into a step the run has already visited: each iteration of a step starts with a new session.
  2. Attempts within an entry continue its session. Every later attempt in the same entry continues the session the entry's first attempt opened. That covers a retry of a retryable harness failure (§6.10), and a re-attempt after output fails result_schema validation (§10.4), such as a reprompt that feeds back the validator's error.
  3. A resume that re-runs a step continues the resumed entry's session. A resume from harness_error (§10.3), schema_violation (§10.4), or budget_exceeded (§10.6) that re-runs the step continues the session of the entry being resumed, even though the re-run counts as a new entry for retry (§6.10). This includes each still-incomplete child re-entered by a resume addressed to a parallel step after a factory-level overrun (§10.6). If the entry being resumed never invoked the harness (for example, budget_exceeded raised on arrival), there is no session to continue, and rule 1 applies.
  4. Other resumes continue nothing. A resume that supplies a result or a routing decision (§11.4) invokes no harness, so no session is opened or continued. A resume from iteration_limit (§10.5) or expression_error (§10.7) re-enters a step whose entry made no harness invocation, so rule 1 applies.

Continuing a session means re-invoking the same harness session, so that the state it holds is kept rather than discarded. An implementation MUST NOT open a new session where these rules require continuing one. If a session that must be continued cannot be, then under rule 3 the resume MUST be rejected (§11.5), and under rule 2 the attempt fails as a non-retryable harness failure (§10.3).

12. Resumability

12.1 Resume equivalence

A run MUST be resumable after process death to a state equivalent to the one it had, where equivalence is defined over FactoryState (§9.2): the same results, the same routing decisions, and the same attempt counts as an uninterrupted run would have produced from the same sequence of StepResults. A run MUST also be resumable after a human-shaped pause of arbitrary duration, with the same equivalence guarantee. last_result (§9.2) at each step follows from the same results and the same routing decisions, so it is covered by this equivalence without a separate guarantee.

This specification requires nothing more of a run's recorded history than this equivalence. What else an implementation records — observability of individual attempts, provenance of a caller-supplied result (§10.4), who owns a run and whether that can change — is a property of the implementation's own operational tooling, not of the SFML file format this specification defines, and this specification imposes no requirement on it. How that data is stored, indexed, or queried is likewise the implementation's own business.

13. Versioning and extensibility

13.1 Version identifier

A factory document declares the version of this specification it targets in its top-level sfml field, an SFMLVersionString (§6.1) such as "v0.1". An implementation MUST reject a document declaring a version it does not support.

13.2 Compatibility policy

A version of this specification MAY add new top-level or step-level fields; it MUST NOT change the meaning of a field already defined by a prior version. Because unknown fields are rejected (§5.4), an implementation of a given version MUST NOT be expected to accept a document that uses a field only a later version defines.

13.4 Extension points

Everything a runtime may vary is confined to two named, declared places: a step's harness reference (§6.12) and its harness_config. A conforming implementation MUST NOT introduce implementation-specific variation anywhere else in the data model; any such variation belongs behind these two fields, where an author can see it named in the file rather than infer it from which implementation happens to be running the factory.


Annex A (normative) — JSON Schema

The JSON Schema for the surface syntax of a factory document is sfml.schema.json in this repository. Where it conflicts with the normative text of this document, this document governs, per §4.4.

Annex B (normative) — Conformance test suite

The conformance test suite is the conformance/ directory of this repository, organized into one top-level category per conformance class (§4.1): parser/ for the Parser rules of §4.3, lint/ for the Linter rules of §4.3, and runner/ for the execution-model behavior of clauses 9–11. Each holds one subdirectory per case, and an implementation conforms with respect to a given class once its test suite passes every case in that class's directory. A Linter need not pass runner/; a Runner, however, MUST pass parser/ and lint/ as well as runner/, per §4.1.3.

A case's directory includes the SFML file (or, for parser/, the raw document) under test, valid or invalid as the case requires, and any prompt files its agent steps name; for a lint/ case, the diagnostic identifier (§8.7) it MUST raise, if any; for a runner/ case, the FactoryState (§9.2) the case MUST produce, and a result file stating whether the case is expected to succeed or to raise a named exception (clause 10); and, where the case invokes an agent step, a script of canned agent messages for the mock harness to play back. A conforming implementation MUST supply a mock harness capable of consuming these scripts and, against them, reproducing the routing trace and FactoryState a case declares. The exact file formats are documented in conformance/README.md, not in this document.

Annex C (informative) — Worked example

The worked example lives in example/, not in this document and not in conformance/: it is maintained separately so it can prioritize being a clear, readable factory over being an exhaustive conformance case. It runs on the mock harness of Annex B. This annex is informative; nothing in example/ is itself normative, though the clauses it illustrates remain normative.