Workflows¶
A workflow in NOMAD is a record of provenance: it describes how inputs were used by tasks to produce outputs. It represents a process that has taken place or is documented in stored data. It is not a workflow engine and does not execute the tasks it describes.
The inputs, tasks, and outputs can refer to archive sections in one entry or across multiple entries. A project groups files and entries for management, access, and publication, whereas a workflow expresses relationships between the data stored in those entries.
Related pages¶
- How-to guides > ... > Create custom workflows
- Explanation > Processing
- How-to guides > Work with schemas > Define a schema
- Tutorials > Use NOMAD as an ELN > Use built-in ELN templates
The built-in abstract workflow schema¶
NOMAD stores the current workflow representation in the top-level workflow2
section of an entry archive. This section contains a Workflow instance from
nomad.datamodel.metainfo.workflow.
%%{init: {"themeVariables": {"fontSize": "18px"}}}%%
classDiagram
direction TB
class EntryArchive
class Workflow
class Task
class TaskReference
class Link
class ArchiveSection
EntryArchive *-- "0..1" Workflow : workflow2
Task <|-- Workflow
Task <|-- TaskReference
Workflow *-- "0..*" Task : tasks
Task *-- "0..*" Link : inputs
Task *-- "0..*" Link : outputs
TaskReference --> Task : task
Task --> ArchiveSection : section
Link --> ArchiveSection : section
Diagram legend
| Notation | Meaning |
|---|---|
<\|-- with an open triangular arrowhead |
Inheritance from a subclass to its parent class |
*-- with a filled diamond |
Composition of contained subsections |
--> with a plain arrowhead |
Reference to another section or task |
0..1 |
Zero or one instance |
0..* |
Zero or more instances |
The referenced ArchiveSection can, for example, belong to run, data, or
another workflow2.
The model has four central components:
- A link represents one input or output. Its
sectionpoints to the archive section containing the data, while itsnameprovides a label for displays such as the workflow graph. - A task represents an activity that used inputs to produce outputs. A
task's optional
sectioncan identify the archive section describing that activity. - A task reference is a proxy for a task or workflow defined elsewhere, such as in another entry. It can supply a local name, inputs, and outputs or obtain them from the referenced task.
- A workflow is a task that also contains other tasks. Because
Workflowinherits fromTask, a workflow can itself be used as a task in another workflow.
The model stores references rather than separate graph edges. When links on different nodes point to the same archive section, NOMAD can represent the shared data as a connection in the workflow graph. References can point to sections in the same entry or to entries elsewhere on the same NOMAD deployment.
Note
Some older entry archives use the legacy top-level workflow section. New
workflow data and schemas should use workflow2.
Nested workflow example¶
A task can represent a process with its own inputs, tasks, and outputs. Because
a Workflow is also a Task, a workflow can be contained directly in a parent
workflow. A TaskReference can instead link to a workflow stored elsewhere.
Consider a parent workflow containing a geometry optimization followed by a single-point workflow for a ground-state calculation:
flowchart TB
input["Input system"]
subgraph parent["Relaxation and ground-state workflow"]
direction TB
subgraph optimization["Geometry optimization"]
direction TB
step0(["Optimization step 0"])
calc0["Calculation 0"]
system1["System 1"]
step1(["Optimization step 1"])
calc1["Calculation 1"]
system2["System 2"]
step2(["Optimization step 2"])
calc2["Calculation 2"]
step0 --> calc0
step0 --> system1
system1 --> step1
step1 --> calc1
step1 --> system2
system2 --> step2
step2 --> calc2
end
step2 --> relaxed["Relaxed system"]
optimization --> relaxed
subgraph single_point["Single-point"]
ground_state(["Ground-state calculation task"])
end
relaxed --> single_point
relaxed --> ground_state
end
input --> optimization
input --> step0
ground_state --> result["Ground-state result"]
single_point --> result
style parent stroke-dasharray: 5 5
style optimization stroke-dasharray: 5 5
style single_point stroke-dasharray: 5 5
Rounded nodes represent tasks, rectangles represent referenced archive sections, dashed boxes represent workflows, and arrows represent links. Each optimization task produces a calculation section and an updated system section that becomes the input of the next task. The input-system link identifies both the geometry optimization's global input and the first task's input. Likewise, the relaxed-system link identifies both the final optimization task's output and the geometry optimization's global output. The relaxed system is then reused as the input of the ground-state calculation task and as the global input of its single-point workflow. Likewise, the ground-state result is both the task output and the global output of the single-point workflow.
The tasks and referenced sections may be stored together or in separate entries without changing these logical relationships. Directly contained and referenced workflows both produce a hierarchical provenance graph while allowing each referenced entry to remain independently accessible.
Custom and standardized workflows¶
The distinction between a custom and a standardized workflow concerns the schema used to describe it, not whether a person or software created the workflow entry.
A custom workflow uses the general Workflow, Task, and Link model
directly. Its author selects the relevant archive sections and defines how they
are connected. This is flexible enough to document processes that do not yet
have a domain-specific workflow schema, while still enabling common graph and
navigation tools.
A standardized workflow uses a specialized Workflow subclass with a
defined scientific meaning and structure. Such a schema can add method and
result sections, workflow-specific quantities, references, and normalization
logic. Examples in NOMAD's simulation schema include SinglePoint,
GeometryOptimization, and GW. Plugins can provide additional specialized
workflow schemas for other domains.
Standardization allows tools to rely on more than the general provenance
graph. A shared schema can support consistent normalization, search,
validation, and domain-specific presentation. The general inputs, tasks, and
outputs remain available because the specialized schema inherits from
Workflow.
How workflow data is created¶
The same workflow model can be populated through several routes:
- Parsers and normalizers can create a workflow while processing supported files. For example, simulation parsers installed on NOMAD Central can create specialized workflows for recognized calculations.
- ELN and other schema normalizers can translate structured records, such as experiment activities and steps, into the general workflow model.
- Archive YAML files can define a custom workflow explicitly or
instantiate an accessible specialized workflow schema with
m_def. - Plugins can define specialized workflow schemas and the parsing or normalization logic that populates them.
These routes can produce either general or specialized workflows. Their availability depends on the parsers, schemas, and plugins installed in a NOMAD deployment.