Skip to content

Workflows

A workflow in NOMAD is a record of provenance: it describes how inputs were used by tasks to produce outputs. It represents a process that has taken place or is documented in stored data. It is not a workflow engine and does not execute the tasks it describes.

The inputs, tasks, and outputs can refer to archive sections in one entry or across multiple entries. A project groups files and entries for management, access, and publication, whereas a workflow expresses relationships between the data stored in those entries.

The built-in abstract workflow schema

NOMAD stores the current workflow representation in the top-level workflow2 section of an entry archive. This section contains a Workflow instance from nomad.datamodel.metainfo.workflow.

%%{init: {"themeVariables": {"fontSize": "18px"}}}%%
classDiagram
    direction TB
    class EntryArchive
    class Workflow
    class Task
    class TaskReference
    class Link
    class ArchiveSection

    EntryArchive *-- "0..1" Workflow : workflow2
    Task <|-- Workflow
    Task <|-- TaskReference
    Workflow *-- "0..*" Task : tasks
    Task *-- "0..*" Link : inputs
    Task *-- "0..*" Link : outputs
    TaskReference --> Task : task
    Task --> ArchiveSection : section
    Link --> ArchiveSection : section
The core classes and relationships used to represent workflow provenance. Framework-level inheritance and other schema details are omitted.

Diagram legend

Notation Meaning
<\|-- with an open triangular arrowhead Inheritance from a subclass to its parent class
*-- with a filled diamond Composition of contained subsections
--> with a plain arrowhead Reference to another section or task
0..1 Zero or one instance
0..* Zero or more instances

The referenced ArchiveSection can, for example, belong to run, data, or another workflow2.

The model has four central components:

  • A link represents one input or output. Its section points to the archive section containing the data, while its name provides a label for displays such as the workflow graph.
  • A task represents an activity that used inputs to produce outputs. A task's optional section can identify the archive section describing that activity.
  • A task reference is a proxy for a task or workflow defined elsewhere, such as in another entry. It can supply a local name, inputs, and outputs or obtain them from the referenced task.
  • A workflow is a task that also contains other tasks. Because Workflow inherits from Task, a workflow can itself be used as a task in another workflow.

The model stores references rather than separate graph edges. When links on different nodes point to the same archive section, NOMAD can represent the shared data as a connection in the workflow graph. References can point to sections in the same entry or to entries elsewhere on the same NOMAD deployment.

Note

Some older entry archives use the legacy top-level workflow section. New workflow data and schemas should use workflow2.

Nested workflow example

A task can represent a process with its own inputs, tasks, and outputs. Because a Workflow is also a Task, a workflow can be contained directly in a parent workflow. A TaskReference can instead link to a workflow stored elsewhere.

Consider a parent workflow containing a geometry optimization followed by a single-point workflow for a ground-state calculation:

flowchart TB
    input["Input system"]

    subgraph parent["Relaxation and ground-state workflow"]
        direction TB
        subgraph optimization["Geometry optimization"]
            direction TB
            step0(["Optimization step 0"])
            calc0["Calculation 0"]
            system1["System 1"]
            step1(["Optimization step 1"])
            calc1["Calculation 1"]
            system2["System 2"]
            step2(["Optimization step 2"])
            calc2["Calculation 2"]

            step0 --> calc0
            step0 --> system1
            system1 --> step1
            step1 --> calc1
            step1 --> system2
            system2 --> step2
            step2 --> calc2
        end

        step2 --> relaxed["Relaxed system"]
        optimization --> relaxed
        subgraph single_point["Single-point"]
            ground_state(["Ground-state calculation task"])
        end

        relaxed --> single_point
        relaxed --> ground_state
    end

    input --> optimization
    input --> step0
    ground_state --> result["Ground-state result"]
    single_point --> result

    style parent stroke-dasharray: 5 5
    style optimization stroke-dasharray: 5 5
    style single_point stroke-dasharray: 5 5
A parent workflow containing nested geometry-optimization and single-point workflows.

Rounded nodes represent tasks, rectangles represent referenced archive sections, dashed boxes represent workflows, and arrows represent links. Each optimization task produces a calculation section and an updated system section that becomes the input of the next task. The input-system link identifies both the geometry optimization's global input and the first task's input. Likewise, the relaxed-system link identifies both the final optimization task's output and the geometry optimization's global output. The relaxed system is then reused as the input of the ground-state calculation task and as the global input of its single-point workflow. Likewise, the ground-state result is both the task output and the global output of the single-point workflow.

The tasks and referenced sections may be stored together or in separate entries without changing these logical relationships. Directly contained and referenced workflows both produce a hierarchical provenance graph while allowing each referenced entry to remain independently accessible.

Custom and standardized workflows

The distinction between a custom and a standardized workflow concerns the schema used to describe it, not whether a person or software created the workflow entry.

A custom workflow uses the general Workflow, Task, and Link model directly. Its author selects the relevant archive sections and defines how they are connected. This is flexible enough to document processes that do not yet have a domain-specific workflow schema, while still enabling common graph and navigation tools.

A standardized workflow uses a specialized Workflow subclass with a defined scientific meaning and structure. Such a schema can add method and result sections, workflow-specific quantities, references, and normalization logic. Examples in NOMAD's simulation schema include SinglePoint, GeometryOptimization, and GW. Plugins can provide additional specialized workflow schemas for other domains.

Standardization allows tools to rely on more than the general provenance graph. A shared schema can support consistent normalization, search, validation, and domain-specific presentation. The general inputs, tasks, and outputs remain available because the specialized schema inherits from Workflow.

How workflow data is created

The same workflow model can be populated through several routes:

  • Parsers and normalizers can create a workflow while processing supported files. For example, simulation parsers installed on NOMAD Central can create specialized workflows for recognized calculations.
  • ELN and other schema normalizers can translate structured records, such as experiment activities and steps, into the general workflow model.
  • Archive YAML files can define a custom workflow explicitly or instantiate an accessible specialized workflow schema with m_def.
  • Plugins can define specialized workflow schemas and the parsing or normalization logic that populates them.

These routes can produce either general or specialized workflows. Their availability depends on the parsers, schemas, and plugins installed in a NOMAD deployment.