Skip to content

How to define a schema

A schema tells NOMAD what your data looks like: which sections it is made of, which quantities each section holds, and how those sections relate to each other. This page covers every concept you need to write one.

Schemas can be written in two syntaxes, Python and YAML, and they describe exactly the same thing. Every example below is shown in both — use the tabs to switch language, and your choice carries down the whole page. If you have not picked a syntax yet, see How-to guides > Work with schemas > Start working with schemas.

The running example

Every example on this page builds the same small schema for a materials lab:

  • a Sample — the specimen being studied;
  • a Process carried out on it, specialized into Evaporation and Annealing;
  • an Instrument that a process was carried out on.
classDiagram
    class Sample {
        name: str
        substrate_type: enum
        tags: str[]
    }
    class Process {
        start_time: Datetime
        instrument: Instrument
    }
    class Evaporation {
        chamber_pressure: float
    }
    class Annealing {
        temperature: float
    }
    class Instrument {
        name: str
    }
    Sample "1" *-- "0..*" Process : processes
    Process <|-- Evaporation
    Process <|-- Annealing
    Process ..> Instrument : references

Define a section

A section groups related data. It is the basic building block of every schema: you define a section once, and NOMAD can then create, store, search and display any number of instances of it.

A section is a Python class. All definitions must sit between the SchemaPackage() constructor and the __init_metainfo__() call that finalizes them.

from nomad.datamodel.data import EntryData
from nomad.metainfo import Quantity, SchemaPackage

m_package = SchemaPackage()


class Instrument(EntryData):
    """A piece of equipment used to carry out a process."""

    name = Quantity(type=str)


m_package.__init_metainfo__()

A section is an entry under definitions.sections. The whole definitions block is the schema package.

definitions:
  name: 'Sample management example schema'
  sections:
    Instrument:
      description: A piece of equipment used to carry out a process.
      base_sections:
        - nomad.datamodel.data.EntryData
      quantities:
        name:
          type: str

Inheriting EntryData is what makes a section usable as the root of an entry — it is why an Instrument can exist as an entry of its own. See Inherit from a base section.

Add quantities

A quantity is a single piece of data: a name, a number, a date, an array. Quantities are where your measurements actually live.

from nomad.metainfo import MEnum, Quantity


class Sample(EntryData):
    """A specimen that a sequence of processes is carried out on."""

    name = Quantity(
        type=str,
        description='A short name for this sample.',
    )
    substrate_type = Quantity(
        type=MEnum('silicon', 'glass', 'sapphire'),
        description='The material the sample was grown on.',
    )
    tags = Quantity(
        type=str,
        shape=['*'],
        description='Free-form labels used to group samples.',
    )
Sample:
  description: A specimen that a sequence of processes is carried out on.
  base_sections:
    - nomad.datamodel.data.EntryData
  quantities:
    name:
      type: str
      description: A short name for this sample.
    substrate_type:
      type:
        type_kind: Enum
        type_data:
          - silicon
          - glass
          - sapphire
      description: The material the sample was grown on.
    tags:
      type: str
      shape: ['*']
      description: Free-form labels used to group samples.

Four attributes carry most of the meaning:

  • type — what values are allowed. Python types (str, int, float, bool), NumPy types (np.float64), Datetime, an enumeration, or another section or quantity to make a reference.
  • shape — the dimensionality. Omit it for a single value, ['*'] for a list, [3, 3] for a 3-by-3 matrix, ['n_atoms', 3] to tie a dimension to another quantity.
  • unit — a physical unit such as pascal or m/s**2. Values are converted to it on assignment.
  • description — what the quantity means. This is shown in the GUI and the Metainfo browser, so it is worth writing.

For the full type table, the shape rules and the naming conventions, see Reference > Schema language.

Add a unit

Quantities that hold a physical measurement should declare a unit. NOMAD then knows what the number means, can convert it, and can display it in whatever unit the reader prefers.

class Annealing(Process):
    temperature = Quantity(
        type=float,
        unit='kelvin',
        description='The temperature the sample was held at.',
    )
Annealing:
  base_section: Process
  quantities:
    temperature:
      type: float
      unit: kelvin
      description: The temperature the sample was held at.

Units are parsed by Pint, so both names and expressions work: m, meter, mm, m/s, m/s**2. See How-to guides > ... > Work with units.

Add subsections

A subsection nests one section inside another, building a containment hierarchy. Use it when the nested data is genuinely part of its parent — a sample's processes belong to that sample and have no independent existence.

Set repeats when the parent can hold many of them.

from nomad.metainfo import SubSection


class Sample(EntryData):
    processes = SubSection(section=Process, repeats=True)
Sample:
  sub_sections:
    processes:
      section: Process
      repeats: true

If the nested data has a life of its own — an instrument used by many processes — use a reference instead of a subsection.

Inherit from a base section

Inheritance lets you build a specialized definition from a more abstract one. The specialization gets every property of its base and can add more, so shared quantities are written once.

Here Evaporation and Annealing both inherit start_time and instrument from Process:

Inheritance is ordinary Python class inheritance.

from nomad.metainfo import Datetime, Quantity


class Process(ArchiveSection):
    """A single step carried out on a sample."""

    start_time = Quantity(type=Datetime)


class Evaporation(Process):
    chamber_pressure = Quantity(type=float, unit='pascal')


class Annealing(Process):
    temperature = Quantity(type=float, unit='kelvin')

Use base_section for a single base, or base_sections for a list.

Process:
  description: A single step carried out on a sample.
  quantities:
    start_time:
      type: Datetime
Evaporation:
  base_section: Process
  quantities:
    chamber_pressure:
      type: float
      unit: pascal
Annealing:
  base_section: Process
  quantities:
    temperature:
      type: float
      unit: kelvin

Sections support multiple inheritance, which is how one section can be both a specialized concept and an entry root section at the same time.

Choose the right base section

  • Inherit from the most specialized base section that still fits your data. The more specific the base, the more built-in NOMAD functionality applies to your entries automatically.
  • Inherit nomad.datamodel.data.EntryData for any section that should be the root of an entry. It is an abstract placeholder, and it is what makes your section appear under an entry's data.
  • Inherit nomad.datamodel.data.ArchiveSection for any section that needs a normalize function. EntryData already includes it.

Commonly used built-in base sections

Section definition or package Purpose
nomad.datamodel.EntryArchive The root object of all NOMAD entries.
nomad.datamodel.EntryMetadata Standard NOMAD metadata: ids, upload, processing and author information.
nomad.datamodel.data.EntryData The abstract section definition for an entry's data section.
nomad.datamodel.data.ArchiveSection Adds support for normalize functions.
nomad.datamodel.metainfo.basesections.* The entity-activity base sections: samples, instruments, processes, measurements, analyses.
nomad.datamodel.metainfo.workflow.* The definitions NOMAD uses to model workflows.
nomad.parsing.tabular.TableData Inherit parsing of .csv and .xls files. See Parse tabular data.
nomad.datamodel.metainfo.basesections.HDF5Normalizer Link quantities to HDF5 datasets for large data. See How-to guides > ... > Handle large data.

For what these mean and why they are shaped the way they are, see Explanation > Base sections. For the generated per-class listing, see Reference > Base sections.

Rely on polymorphism

A subsection or reference declares the most abstract section it accepts, and any specialization of that section is then a valid value. Because Sample.processes declares Process, a sample can hold an Evaporation and an Annealing in the same list, and each keeps its own extra quantities.

This is what lets one schema define a relationship and another schema extend what can fill it.

Iterating a subsection yields instances of whichever subclass was actually stored.

sample = Sample(name='S1')
sample.processes.append(Evaporation(chamber_pressure=1e-5))
sample.processes.append(Annealing(temperature=600))

for process in sample.processes:
    print(type(process).__name__, process.start_time)

# The specialized quantities are available on each instance:
sample.processes[0].chamber_pressure
sample.processes[1].temperature

Each item in the list carries an m_def naming which specialization it is.

data:
  m_def: Sample
  name: S1
  processes:
    - m_def: Evaporation
      chamber_pressure: 1.0e-05
    - m_def: Annealing
      temperature: 600

Note

m_def is only needed when the section definition cannot be worked out from context. A subsection that accepts exactly one type does not need it; a polymorphic one does.

Subsections express a part-of relationship. When data is linked rather than contained — a process pointing at the instrument it ran on, where that instrument is shared by many processes — use a reference.

A reference is a uni-directional link from a source quantity to a target. The target can be a whole section, or a single quantity inside another section, and what you put in type decides which. The two kinds differ in what you assign and what you get back:

Reference kind type is You assign You read back
Section reference a section definition the target section the target section
Quantity reference a quantity definition the section holding the quantity that quantity's value

Either kind can hold many targets: give it a shape of ['*']. The shape describes the number of references, not the shape of the data they point at.

Reference a section

Use the target's section definition as the type:

class Process(ArchiveSection):
    instrument = Quantity(
        type=Instrument,
        description='The instrument used for this process.',
    )
Process:
  quantities:
    instrument:
      type: Instrument
      description: The instrument used for this process.

In memory the quantity holds the target section. When the archive is saved, it is serialized as a URL — a path from the archive root such as #/data/processes/0, or a longer form that crosses into another entry or another NOMAD installation. The full list of reference forms is in Reference > Schema language > Reference forms.

Reference a quantity

Use a single quantity as the type when the source needs one value from elsewhere in the archive rather than the whole section. Reading the quantity gives you that value, so the data is exposed without being copied.

class Process(ArchiveSection):
    instrument_name = Quantity(
        type=Instrument.name,
        description='The name of the instrument used for this process.',
    )
Process:
  quantities:
    instrument_name:
      type:
        type_kind: quantity_reference
        type_data: Instrument/name
      description: The name of the instrument used for this process.

Note

Writing the target as a plain string, type: Instrument/name, does not work: a bare string is always read as a section reference. The type_kind form is the only one that makes a quantity reference, and it can only name a quantity defined in the same file.

Note

A quantity reference stores a link to the section that holds the target quantity, never a copy of the value. So you assign the section and read the value:

process.instrument_name = instrument  # assign the section
process.instrument_name  # 'Evaporator A' — read the value

It is serialized as that section's URL with the quantity name appended, for example #/data/instruments/0/name. If the target quantity is not set on the section you assigned, the reference counts as unset and is left out of the archive.

Reference across entries

Because Instrument inherits EntryData, each instrument is its own entry, and a sample in one entry references an instrument in another. In YAML you write that target as a path to the other file:

data:
  m_def: Sample
  name: S1
  processes:
    - m_def: Evaporation
      instrument: ../upload/raw/evaporator.archive.yaml#/data

The declaration of instrument is identical in both languages — only the serialized value differs, and NOMAD resolves it for you when the archive is read.

Note

References are resolved lazily. On loading, a reference becomes a placeholder that is replaced by the real section the first time you access it. This is what makes references between entries, and even between NOMAD installations, possible. See Advanced schema concepts.

Annotate for the GUI

A schema says what data is. Annotations say what NOMAD should do with it — which editor to show, how to plot it, which unit to display. Adding ELN annotations is also what turns a schema into an electronic lab notebook that users can fill in through the browser.

The component you choose has to suit the quantity's type: a string gets a text field, a datetime gets a date picker, an enumeration gets a dropdown, a reference gets a search-and-select field.

Annotations are keyword arguments beginning with a_.

from nomad.datamodel.metainfo.annotations import (
    ELNAnnotation,
    QuantityDisplayAnnotation,
)


class Sample(EntryData):
    name = Quantity(
        type=str,
        a_eln=ELNAnnotation(component='StringEditQuantity'),
    )
    substrate_type = Quantity(
        type=MEnum('silicon', 'glass', 'sapphire'),
        a_eln=ELNAnnotation(component='EnumEditQuantity'),
    )
    thickness = Quantity(
        type=float,
        unit='meter',
        a_eln=ELNAnnotation(component='NumberEditQuantity'),
        a_display=QuantityDisplayAnnotation(unit='nm'),
    )

Annotations are named blocks under m_annotations.

Sample:
  quantities:
    name:
      type: str
      m_annotations:
        eln:
          component: StringEditQuantity
    substrate_type:
      type:
        type_kind: Enum
        type_data: [silicon, glass, sapphire]
      m_annotations:
        eln:
          component: EnumEditQuantity
    thickness:
      type: float
      unit: meter
      m_annotations:
        eln:
          component: NumberEditQuantity
        display:
          unit: nm

The display annotation on thickness sets the unit the GUI shows the value in. It does not change how the value is stored: that is always the declared unit, and a thickness entered in nanometres is converted to metres when saved. The display unit has to be compatible with the declared one.

Annotate a section

Annotations on the section describe the section as a whole rather than one of its properties. The display annotation is the one to reach for first: it decides which properties the GUI shows, and in which order.

Annealing inherits start_time and instrument from Process. A lab that always anneals in the same furnace has no use for the instrument field, so visible drops it and order puts the quantity that matters first.

Section annotations go on a m_def = Section(...) assignment inside the class.

from nomad.datamodel.metainfo.annotations import Filter, SectionDisplayAnnotation
from nomad.metainfo import Section


class Annealing(Process):
    m_def = Section(
        a_display=SectionDisplayAnnotation(
            visible=Filter(exclude=['instrument']),
            order=['temperature', 'start_time'],
        ),
    )
    temperature = Quantity(type=float, unit='kelvin')

Section annotations go in an m_annotations block directly under the section, beside quantities.

Annealing:
  base_section: Process
  m_annotations:
    display:
      visible:
        exclude: [instrument]
      order: [temperature, start_time]
  quantities:
    temperature:
      type: float
      unit: kelvin

visible takes an include list, an exclude list, or both: without include every property of the section starts out visible, and exclude is subtracted afterwards, so a name given in both is excluded. editable is a filter of the same kind, but it renders properties read-only instead of hiding them — useful when a quantity inherited from a base section should be shown but not changed. order lists the properties that come first; everything else follows in declaration order.

Beyond eln and display, annotations control plotting and HDF5 visualization. Every annotation and its arguments is listed in Reference > Annotations.

Populate data

With the schema defined, you create instances of it and fill them in.

Section instances behave like ordinary Python objects. Values are converted to the declared type and unit on assignment.

sample = Sample(name='S1', substrate_type='glass')
sample.tags = ['internal']

evaporation = Evaporation(chamber_pressure=1e-5)
sample.processes.append(evaporation)

print(sample.m_to_json(indent=2))

Methods beginning with m_ provide the metainfo behaviour: m_to_json and m_to_dict serialize, m_from_dict reads back. To set and read properties by name at runtime, see Advanced schema concepts.

Data goes under the top-level data key, with m_def naming the section it instantiates. It can live in the same file as the schema, or in a file of its own.

data:
  m_def: Sample
  name: S1
  substrate_type: glass
  tags: [internal]
  processes:
    - m_def: Evaporation
      chamber_pressure: 1.0e-05

Add normalize functions

A normalize function runs every time an entry is processed, which happens whenever a file is uploaded or changed. Use it to derive values, fill in defaults, or copy data into a more interoperable part of the archive.

Python only

Normalize functions are the one capability YAML schemas do not have, because they are code rather than data. If you need derived values, write the schema in Python. This is the most common reason to move a working YAML schema into a plugin.

In order for the normalize function to tbe triggered, the section must inherit ArchiveSection (EntryData already does):

class Sample(EntryData):
    name = Quantity(type=str)
    sample_id = Quantity(type=str)

    processes = SubSection(section=Process, repeats=True)

    def normalize(self, archive, logger):
        super().normalize(archive, logger)

        if self.sample_id is None and self.name is not None:
            self.sample_id = f'{self.name}--{len(self.processes)}'

You usually want to call super().normalize(...) so multiple inheritance keeps working.

Normalize functions run for every subsection before their parent. To control the order among sections at the same level, set normalizer_level; it defaults to 0 and sections run from low to high.

Usually it is good to design a normalize function so it only needs data from its own section. Use m_parent and m_root to read from the surrounding archive when you must, but avoid writing outside your own section.

Note

A normalize function is not the same thing as a normalizer. A normalizer is a separate plugin entry point that operates on a whole entry. See How-to guides > ... > Normalizers and Explanation > From files to data.

The complete example

The running example, as one working schema. The two files define exactly the same sections and quantities; only the normalize function is Python-only.

from nomad.datamodel.data import ArchiveSection, EntryData
from nomad.datamodel.metainfo.annotations import (
    ELNAnnotation,
    Filter,
    QuantityDisplayAnnotation,
    SectionDisplayAnnotation,
)
from nomad.metainfo import Datetime, MEnum, Quantity, SchemaPackage, Section, SubSection

m_package = SchemaPackage()


class Instrument(EntryData):
    """A piece of equipment used to carry out a process."""

    name = Quantity(
        type=str,
        description='The name this instrument is known by in the lab.',
        a_eln=ELNAnnotation(component='StringEditQuantity'),
    )


class Process(ArchiveSection):
    """A single step carried out on a sample."""

    start_time = Quantity(
        type=Datetime,
        description='When the process was started.',
        a_eln=ELNAnnotation(component='DateTimeEditQuantity'),
    )
    instrument = Quantity(
        type=Instrument,
        description='The instrument used for this process.',
        a_eln=ELNAnnotation(component='ReferenceEditQuantity'),
    )


class Evaporation(Process):
    """Depositing a material onto the sample from the vapour phase."""

    chamber_pressure = Quantity(
        type=float,
        unit='pascal',
        description='The pressure in the chamber during evaporation.',
        a_eln=ELNAnnotation(component='NumberEditQuantity'),
    )


class Annealing(Process):
    """Heating the sample to change its structure."""

    # This section is always carried out in the same furnace, so the inherited
    # "instrument" quantity is hidden and "temperature" is shown first.
    m_def = Section(
        a_display=SectionDisplayAnnotation(
            visible=Filter(exclude=['instrument']),
            order=['temperature', 'start_time'],
        ),
    )
    temperature = Quantity(
        type=float,
        unit='kelvin',
        description='The temperature the sample was held at.',
        a_eln=ELNAnnotation(component='NumberEditQuantity'),
    )


class Sample(EntryData):
    """A specimen that a sequence of processes is carried out on."""

    name = Quantity(
        type=str,
        description='A short name for this sample.',
        a_eln=ELNAnnotation(component='StringEditQuantity'),
    )
    substrate_type = Quantity(
        type=MEnum('silicon', 'glass', 'sapphire'),
        description='The material the sample was grown on.',
        a_eln=ELNAnnotation(component='EnumEditQuantity'),
    )
    thickness = Quantity(
        type=float,
        unit='meter',
        description='The thickness of the sample.',
        a_eln=ELNAnnotation(component='NumberEditQuantity'),
        a_display=QuantityDisplayAnnotation(unit='nm'),
    )
    tags = Quantity(
        type=str,
        shape=['*'],
        description='Free-form labels used to group samples.',
    )
    sample_id = Quantity(
        type=str,
        description='An identifier derived from the name and the number of processes.',
    )

    processes = SubSection(section=Process, repeats=True)

    def normalize(self, archive, logger):
        super().normalize(archive, logger)

        if self.sample_id is None and self.name is not None:
            self.sample_id = f'{self.name}--{len(self.processes)}'


m_package.__init_metainfo__()

# The same schema as examples/schemas/sample_schema.py, written in YAML.
# All definitions live under the top-level "definitions" key, which NOMAD
# interprets as a schema package.
definitions:
  name: 'Sample management example schema'
  sections:
    Instrument:
      description: A piece of equipment used to carry out a process.
      base_sections:
        - nomad.datamodel.data.EntryData
      quantities:
        name:
          type: str
          description: The name this instrument is known by in the lab.
          m_annotations:
            eln:
              component: StringEditQuantity
    Process:
      description: A single step carried out on a sample.
      quantities:
        start_time:
          type: Datetime
          description: When the process was started.
          m_annotations:
            eln:
              component: DateTimeEditQuantity
        instrument:
          type: Instrument
          description: The instrument used for this process.
          m_annotations:
            eln:
              component: ReferenceEditQuantity
    Evaporation:
      description: Depositing a material onto the sample from the vapour phase.
      base_section: Process
      quantities:
        chamber_pressure:
          type: float
          unit: pascal
          description: The pressure in the chamber during evaporation.
          m_annotations:
            eln:
              component: NumberEditQuantity
    Annealing:
      description: Heating the sample to change its structure.
      base_section: Process
      # This section is always carried out in the same furnace, so the inherited
      # "instrument" quantity is hidden and "temperature" is shown first.
      m_annotations:
        display:
          visible:
            exclude: [instrument]
          order: [temperature, start_time]
      quantities:
        temperature:
          type: float
          unit: kelvin
          description: The temperature the sample was held at.
          m_annotations:
            eln:
              component: NumberEditQuantity
    Sample:
      description: A specimen that a sequence of processes is carried out on.
      base_sections:
        - nomad.datamodel.data.EntryData
      quantities:
        name:
          type: str
          description: A short name for this sample.
          m_annotations:
            eln:
              component: StringEditQuantity
        substrate_type:
          type:
            type_kind: Enum
            type_data:
              - silicon
              - glass
              - sapphire
          description: The material the sample was grown on.
          m_annotations:
            eln:
              component: EnumEditQuantity
        thickness:
          type: float
          unit: meter
          description: The thickness of the sample.
          m_annotations:
            eln:
              component: NumberEditQuantity
            display:
              unit: nm
        tags:
          type: str
          shape: ['*']
          description: Free-form labels used to group samples.
        sample_id:
          type: str
          description: An identifier derived from the name and the number of processes.
      sub_sections:
        processes:
          section: Process
          repeats: true