What an annotator declares

An annotator’s output is a set of attributes. Three objects describe an attribute, one per stage of its life, and the distinction between them is what makes the rest of the chapter readable:

  • An AttributeSpec is the static declaration: what the annotator can produce. It is written by the annotator’s author, keyed by source, and does not depend on any pipeline.

  • An AttributeConfig is the attribute as written in pipeline YAML: a source, the name to emit it under, whether it is internal, which aggregator to reduce it with, and any parameters. It is produced by the configuration parser and an extender never constructs one.

  • An Attribute is the runtime attribute a configured annotator exposes through its attributes: the config resolved against the spec, carrying both.

The specification

Annotator.get_attribute_specs returns every attribute the annotator can produce, as a mapping from source to AttributeSpec. It is the catalogue the configuration is checked against — a configured attribute whose source is not a key here is refused — and it is independent of the configuration and of open, so it may use only what the subclass set before it delegated to the base constructor.

A spec carries the source, a value_type, a description shown in the generated pipeline documentation, and four flags:

  • is_default — whether the attribute is emitted when the YAML names no attributes at all. With no attributes: block, every default spec stands in, under its source name.

  • internal_default — whether the attribute is internal unless the YAML says otherwise: computed and available to later annotators through the context, but not written to the output.

  • supports_aggregation — whether an aggregator may be configured for it.

  • attribute_type — "attribute" for a value, or "annotatable" for an attribute that is itself an annotatable and can be named as another annotator’s input_annotatable. An annotatable attribute never supports aggregation.

The value_type is a string naming the type of a single value: int, float, str, object (a value that is not one of the scalar types), list and annotatable are the spellings in use. It is what an aggregator is validated against — a numeric aggregator over a str attribute is refused when the pipeline is built — and what the web annotation interface reads to present a value; the tabular and VCF writers do not consult it.

Defaults per attribute

Some annotators have defaults that do not belong in the spec: a score resource declares, in its own configuration, the aggregator each score should be reduced with. AnnotatorBase.get_attribute_defaults is the hook for those. The base constructor consults it once per attribute: an aggregator key becomes the aggregator when the YAML names none, and every other key becomes an attribute parameter that the YAML’s own parameters override. The default returns an empty mapping, which is right for any annotator whose defaults are already in its specs.

The runtime attribute

Once configured, an annotator’s attributes is a list of Attribute, in configuration order. Each one holds its name (the output column), its source, the resolved internal flag and aggregator, its parameters, and the spec it was resolved against. Two things about it matter to an implementer:

The name is read at answer time, never cached. A pipeline that names one attribute twice renames the later ones, and it does so after every annotator has been constructed. An annotator that captured attr.name while building its queries would answer under names the pipeline has since moved away from. Whatever a query list caches, it must not cache names — walk self.attributes again when you answer. The helpers on AnnotatorBase all do this for you.

Parameters are usage-monitored. An attribute’s parameters (and the annotator’s own, on its AnnotatorInfo) are a mapping that records every key read. After an annotator is built, the pipeline refuses any key that was configured but never read, on the grounds that a parameter nothing consulted is a typo. So an annotator reads every parameter it accepts in its constructor, not lazily on the first annotate.

Aggregation

An aggregator reduces many values to one — a position score read over a region answers one value per position, and max or mean folds them. The attribute names the aggregator; Attribute.fold is the one statement of how it reduces. Whether there is anything to reduce is the annotator’s decision, because the container differs by annotator: the base does not aggregate on an annotator’s behalf, and Writing an annotator says which helper to reach for.

API

class gain.annotation.annotation_pipeline.AttributeSpec(source: str, value_type: str, description: str, is_default: bool = True, internal_default: bool = False, supports_aggregation: bool = True, attribute_type: str = 'attribute')[source]

Describes a single attribute an annotator can produce.

as_dict() → dict[str, Any][source]

Serialize to a response dict.

class gain.annotation.annotation_config.Attribute(name: str, source: str, internal: bool | None = None, aggregator: AggregatorSource | None = None, parameters: ParamsUsageMonitor = <factory>, spec: AttributeSpec | None = None, _documentation: str | None = None)[source]

Runtime attribute instance produced by an annotator.

fold(values: list[Any]) → Any[source]

Reduce values with the aggregator this attribute NAMES.

The one statement of HOW an attribute reduces, kept beside the name it reduces by. The caller decides WHETHER there is anything to reduce, because the container differs by annotator – a list of a score’s values, a mapping of per-gene values – and must hold an aggregator before asking.

A fresh accumulator per call, never a held one: an aggregator is mutable state, and one built here cannot outlive the fold it was built for. Building costs ~0.26 us against the ~0.06 us of clearing a held instance (measured, gain#1133) – a fifth of a microsecond per folded attribute per variant, paid only where a fold actually happens. The name resolution itself is memoised (gain#1157), so what is left is the object, not the parsing.

get_value_type(*, aggregated: bool = True) → str[source]

Value type produced by this attribute.

Pass aggregated=True (default) when the aggregator is known to have run; the aggregator’s output_value_type then takes precedence over the spec’s declared type. Pass aggregated=False when aggregation was skipped (e.g. a scalar value that bypassed a list aggregator) so that the spec type is returned instead. The raw spec type is always accessible via self.spec.value_type.

The type is read off the aggregator’s NAME – the only thing the attribute holds since gain#1133 – through Aggregator.resolve_class(), which is class-level and builds no accumulator.