What an annotator declares
An annotator’s output is a set of attributes. Three objects describe an attribute, one per stage of its life, and the distinction between them is what makes the rest of the chapter readable:
An
AttributeSpecis the static declaration: what the annotator can produce. It is written by the annotator’s author, keyed by source, and does not depend on any pipeline.An
AttributeConfigis the attribute as written in pipeline YAML: a source, the name to emit it under, whether it is internal, which aggregator to reduce it with, and any parameters. It is produced by the configuration parser and an extender never constructs one.An
Attributeis the runtime attribute a configured annotator exposes through itsattributes: the config resolved against the spec, carrying both.
The specification
Annotator.get_attribute_specs
returns every attribute the annotator can produce, as a mapping from source
to AttributeSpec. It is the
catalogue the configuration is checked against — a configured attribute
whose source is not a key here is refused — and it is independent of the
configuration and of open, so it may use only what the subclass set
before it delegated to the base constructor.
A spec carries the source, a value_type, a description shown in
the generated pipeline documentation, and four flags:
is_default— whether the attribute is emitted when the YAML names no attributes at all. With noattributes:block, every default spec stands in, under its source name.internal_default— whether the attribute is internal unless the YAML says otherwise: computed and available to later annotators through the context, but not written to the output.supports_aggregation— whether an aggregator may be configured for it.attribute_type—"attribute"for a value, or"annotatable"for an attribute that is itself an annotatable and can be named as another annotator’sinput_annotatable. An annotatable attribute never supports aggregation.
The value_type is a string naming the type of a single value: int,
float, str, object (a value that is not one of the scalar
types), list and annotatable are the spellings in use. It is what an
aggregator is validated against — a numeric aggregator over a str
attribute is refused when the pipeline is built — and what the web
annotation interface reads to present a value; the tabular and VCF writers
do not consult it.
Defaults per attribute
Some annotators have defaults that do not belong in the spec: a score
resource declares, in its own configuration, the aggregator each score
should be reduced with. AnnotatorBase.get_attribute_defaults
is the hook for those. The base constructor consults it once per attribute:
an aggregator key becomes the aggregator when the YAML names none, and
every other key becomes an attribute parameter that the YAML’s own
parameters override. The default returns an empty mapping, which is right
for any annotator whose defaults are already in its specs.
The runtime attribute
Once configured, an annotator’s
attributes is a list
of Attribute, in configuration
order. Each one holds its name (the output column), its source, the
resolved internal flag and aggregator, its parameters, and the
spec it was resolved against. Two things about it matter to an
implementer:
The name is read at answer time, never cached. A pipeline that names
one attribute twice renames the later ones, and it does so after every
annotator has been constructed. An annotator that captured attr.name
while building its queries would answer under names the pipeline has since
moved away from. Whatever a query list caches, it must not cache names —
walk self.attributes again when you answer. The helpers on
AnnotatorBase all do this for you.
Parameters are usage-monitored. An attribute’s parameters (and the
annotator’s own, on its AnnotatorInfo)
are a mapping that records every key read. After an annotator is built, the
pipeline refuses any key that was configured but never read, on the grounds
that a parameter nothing consulted is a typo. So an annotator reads every
parameter it accepts in its constructor, not lazily on the first annotate.
Aggregation
An aggregator reduces many values to one — a position score read over a
region answers one value per position, and max or mean folds them.
The attribute names the aggregator; Attribute.fold
is the one statement of how it reduces. Whether there is anything to reduce
is the annotator’s decision, because the container differs by annotator: the
base does not aggregate on an annotator’s behalf, and
Writing an annotator says which helper to reach for.
API
- class gain.annotation.annotation_pipeline.AttributeSpec(source: str, value_type: str, description: str, is_default: bool = True, internal_default: bool = False, supports_aggregation: bool = True, attribute_type: str = 'attribute')[source]
Describes a single attribute an annotator can produce.
- class gain.annotation.annotation_config.Attribute(name: str, source: str, internal: bool | None = None, aggregator: AggregatorSource | None = None, parameters: ParamsUsageMonitor = <factory>, spec: AttributeSpec | None = None, _documentation: str | None = None)[source]
Runtime attribute instance produced by an annotator.
- fold(values: list[Any]) Any[source]
Reduce
valueswith the aggregator this attribute NAMES.The one statement of HOW an attribute reduces, kept beside the name it reduces by. The caller decides WHETHER there is anything to reduce, because the container differs by annotator – a list of a score’s values, a mapping of per-gene values – and must hold an aggregator before asking.
A fresh accumulator per call, never a held one: an aggregator is mutable state, and one built here cannot outlive the fold it was built for. Building costs ~0.26 us against the ~0.06 us of clearing a held instance (measured, gain#1133) – a fifth of a microsecond per folded attribute per variant, paid only where a fold actually happens. The name resolution itself is memoised (gain#1157), so what is left is the object, not the parsing.
- get_value_type(*, aggregated: bool = True) str[source]
Value type produced by this attribute.
Pass
aggregated=True(default) when the aggregator is known to have run; the aggregator’soutput_value_typethen takes precedence over the spec’s declared type. Passaggregated=Falsewhen aggregation was skipped (e.g. a scalar value that bypassed a list aggregator) so that the spec type is returned instead. The raw spec type is always accessible viaself.spec.value_type.The type is read off the aggregator’s NAME – the only thing the attribute holds since gain#1133 – through
Aggregator.resolve_class(), which is class-level and builds no accumulator.