Writing an annotator

Two classes define what an annotator is. Annotator is the interface the pipeline drives: open, annotate, close, and the attributes it produces. AnnotatorBase is the class an annotator actually extends: it implements the interface once, on top of two narrower hooks that a subclass fills in. Every annotator GAIn ships, and every in-tree plugin, is built on AnnotatorBase; the interface class is documented so you can read the contract the base keeps, not so you can implement it yourself.

What the base does for you

The base constructor takes the pipeline and an AnnotatorInfo — the annotator’s type, id, configured attributes, parameters and resources, as parsed from the YAML. It checks each configured attribute against get_attribute_specs(), resolves its name, aggregator and parameters (consulting get_attribute_defaults()), and requires a work_dir parameter — a directory of the annotator’s own, which the pipeline injects into every annotator’s parameters before the factory runs and which open() creates.

Because get_attribute_specs is called from the base constructor, a subclass that computes its specs from something — a resource’s score list, say — sets that something before it calls super().__init__.

Two YAML parameters are handled by the pipeline rather than by the annotator, and a subclass never sees them: input_annotatable, which substitutes an annotatable produced earlier in the pipeline for the input row’s own, and an attribute’s value_transform, which rewrites the answer on its way into the context. Both are applied by wrapping the built annotator, so the contract below is stated for the annotator alone.

The contract

A subclass implements two members, and may override a handful more.

Always:

  • get_attribute_specs() — every attribute the annotator can produce, keyed by source (What an annotator declares).

  • _do_annotate() — produce the attributes for one annotatable. It receives the annotatable (never None; the base answers the empty result for that before delegating) and the context, and answers an AnnotatedValues: a mapping keyed by attribute name, every value finished. The base hands it back to the pipeline as it is; nothing is reduced or renamed on the annotator’s behalf.

When there is a reason:

  • _do_batch_annotate() — the same contract for a sequence of annotatables: one AnnotatedValues per input, in input order, the empty result for each None. The default loops _do_annotate and is correct for every annotator; override it only when the backend has a genuinely batched path — an external tool run once over a file, a model that predicts a batch at a time. A batch-only annotator, one with no per-item path, overrides annotate to raise NotImplementedError as well; the demo_annotator adapters and the VEP annotators are the shipped examples.

  • open() and close() — when the annotator holds a resource. Overrides open what they query, call the base and return self; close must be safe on an annotator never opened and safe twice.

  • get_attribute_defaults() — when per-attribute defaults live somewhere other than the spec.

  • used_context_attributes — when the annotator reads attributes an earlier annotator produced. Name them here: the pipeline builds its dependency graph from this tuple, and a reannotation reruns the annotator when a named attribute’s producer changes. Reading the context without declaring it works for a plain annotation run, and silently goes stale under reannotation.

  • ACCEPTED_RESOURCE_TYPES and resolve_resource() — when the annotator takes a resource_id naming a typed genomic resource. Declare the types it accepts, resolve through the classmethod, and the refusal for a resource of the wrong type is phrased the same way as every other annotator’s.

Never: annotate() and batch_annotate(). They are the pipeline’s side of the contract, and the base already routes them to the two hooks. The one exception is the batch-only annotate above.

Shaping the answer

Every AnnotatedValues is keyed by attribute name, and names must be read when the annotator answers, never cached beside the queries it built — What an annotator declares says why. The base provides four helpers that read self.attributes at answer time so a subclass does not have to:

  • _every() — one value under every attribute’s name. For an annotator with one thing to say however many attributes expose it: a lifted-over annotatable, a renamed chromosome, a yes/no decision.

  • _from_sources() — a source-keyed mapping answered by attribute name, nothing folded. For values that are final as they stand — a point read’s one value per score, a tool’s output row. An absent source answers None.

  • fold_own_values() — the folding twin of _from_sources, for an annotator whose values are its own rather than a score’s record stream: each attribute’s source value is reduced by the aggregator the attribute names, if it names one and the value is a list, and answered under the attribute’s name.

  • _pair_aggregated() and _pair_all() — for a score annotator whose score did the folding during the read, and answered one value per query in attribute order. They pair those values back onto the attributes by position, and refuse a read that answered the wrong number of them.

_empty_result() is _every(None): the answer for a None annotatable, and what an annotator returns when a guard fires — a chromosome the resource does not have, a region past a length cutoff.

A minimal annotator

The smallest complete annotator declares one attribute, reads the context, and answers through _every. This is the shape of the plugin GAIn Python interface builds in its fifth section, whose full adapter is downloadable there:

from typing import Any

from gain.annotation.annotatable import Annotatable
from gain.annotation.annotation_pipeline import (
    AnnotationPipeline,
    Annotator,
    AnnotatorInfo,
    AttributeSpec,
)
from gain.annotation.annotator_base import AnnotatedValues, AnnotatorBase


class FollowupAnnotator(AnnotatorBase):
    """Flag a variant for follow-up from attributes already computed."""

    def get_attribute_specs(self) -> dict[str, AttributeSpec]:
        return {
            "followup": AttributeSpec(
                source="followup",
                value_type="str",
                description="Whether the variant is selected for follow-up",
            ),
        }

    @property
    def used_context_attributes(self) -> tuple[str, ...]:
        return ("phyloP7way", "clinical_significance")

    def _do_annotate(
        self, annotatable: Annotatable, context: dict[str, Any],
    ) -> AnnotatedValues:
        conserved = (context.get("phyloP7way") or 0) > 0
        pathogenic = context.get("clinical_significance") == "Pathogenic"
        return self._every("yes" if conserved and pathogenic else "no")


def build_followup_annotator(
    pipeline: AnnotationPipeline, info: AnnotatorInfo,
) -> Annotator:
    return FollowupAnnotator(pipeline, info)

The module-level function at the end is the factory: the callable a pipeline resolves the annotator’s type to, taking the pipeline and the AnnotatorInfo and returning a built annotator. Annotator plugins says how it is registered under a name.

Annotators that wrap a tool

An annotator that runs an external program — in a subprocess, or in a container — is a batch annotator by nature: it writes its inputs to a file or a stream, runs the tool once, and reads the answers back. The three in-tree plugin packages are the worked examples, and are cited here rather than paraphrased:

  • demo_annotator — four annotators, purpose-built as an example. Its adapter module shows the adapter pattern twice over the same toy tool: once through temporary files in the annotator’s work_dir, once over the tool’s stdin and stdout. Its two demo_annotate_*_adapter modules show an adapter that hands a GRR resource — gene models, a reference genome — to a tool that reads files, through a caching repository rooted in the work_dir.

  • vep_annotator — two real annotators over Ensembl VEP in a container, on a shared base that builds VEP’s input and parses its output.

  • spliceai_annotator — a real annotator over a model, with a per-item _do_annotate and a batched _do_batch_annotate that predicts many requests at once; the one in-tree annotator with both paths.

DockerAnnotator is the base the containerised ones share: it holds a Docker client, prepares the images the annotator needs in open, and leaves the run of one batch abstract.

API

class gain.annotation.annotation_pipeline.Annotator(pipeline: AnnotationPipeline | None, info: AnnotatorInfo)[source]

An annotator produces a set of attributes for a given annotatable.

The pipeline drives the lifecycle: open() before the first annotate(), close() once at the end. An annotator may assume it is open when asked to annotate, and does not open itself on demand. Implementations extend AnnotatorBase, which handles configuration and the None annotatable, rather than this class directly.

abstractmethod annotate(annotatable: Annotatable | None, context: dict[str, Any]) → dict[str, Any][source]

Produce this annotator’s attributes for one annotatable.

Returns a mapping from attribute name (not source) to value, with every attribute in attributes present. annotatable is None when the input row has none – an unparsable variant, a liftover that found nothing – and the answer is then every attribute set to None, never an exception. context holds the attributes of the annotators before this one: read what used_context_attributes declares and do not write to it – the pipeline merges the returned mapping into it. May assume open() has run. An annotator that only works in batches raises NotImplementedError here and overrides batch_annotate().

abstract property attributes: list[Attribute]

The attributes this annotator produces, in output order.

Configured attributes: names, sources, aggregators and parameters already resolved against get_attribute_specs().

batch_annotate(annotatables: Sequence[Annotatable | None], contexts: list[dict[str, Any]], batch_work_dir: str | None = None) → Iterable[dict[str, Any]][source]

Annotate many annotatables: one result per input, in order.

The default calls annotate() once per pair, lazily, and is correct for every annotator. Override it only when the backend has a genuinely batched path – an external tool run once over a file, say – and keep the same contract: exactly one result per annotatable, in input order; the empty result for a None annotatable; contexts read, not written. batch_work_dir is a scratch directory the caller may offer, None when it does not; the default ignores it.

close() → None[source]

Release what open() acquired and mark the annotator closed.

Safe on an annotator never opened, and safe twice; overrides keep it so and call the base. The pipeline calls it once per annotator and logs, rather than propagates, what it raises.

abstractmethod get_attribute_specs() → dict[str, AttributeSpec][source]

Every attribute this annotator can produce, keyed by source.

The catalogue the configuration is checked against: a configured attribute whose source is not a key here is refused. Independent of the configuration and of open(). AnnotatorBase calls it from its constructor, so it may use only what the subclass set before delegating there.

get_info() → AnnotatorInfo[source]

The annotator info this annotator was built from.

An AnnotatorInfo: its type, id, configured attributes, parameters and resources.

is_open() → bool[source]

Whether open() has run and close() has not since.

open() → Annotator[source]

Acquire resources and mark the annotator open; returns self.

The base only flips the flag. Overrides open the resources they query, call the base and return self. Opening an already-open annotator must be harmless.

property resource_ids: set[str]

The ids of resources, as a set.

property resources: list[GenomicResource]

The genomic resources this annotator was configured with.

property used_context_attributes: tuple[str, ...]

Names of upstream attributes this annotator reads from context.

Empty by default. An annotator that reads an attribute another annotator produced – a gene list, say – names it here: the pipeline builds its dependency graph from this tuple, and a reannotation reruns this annotator when a named attribute’s producer changes. Every name must be an attribute of an earlier annotator in the same pipeline.

class gain.annotation.annotator_base.AnnotatorBase(pipeline: AnnotationPipeline | None, info: AnnotatorInfo)[source]

Base implementation of the Annotator class.

The class every in-tree annotator extends. Its constructor checks the configured attributes against get_attribute_specs(), resolves each one’s name, aggregator and parameters (consulting get_attribute_defaults()), and requires a work_dir parameter. A subclass implements get_attribute_specs and _do_annotate(); overrides _do_batch_annotate() when it has a batched path; and overrides get_attribute_defaults(), open() and close() when it has defaults or resources. annotate() and batch_annotate() are left alone, except by a batch-only annotator, which makes annotate() refuse.

ACCEPTED_RESOURCE_TYPES: ClassVar[tuple[str, ...]] = ()

The resource types this annotator’s resource_id may name.

An annotator that consumes a typed genomic resource states them here, and resolves its resource through resolve_resource(). Before gain#1329 the same fact was written once per annotator in whatever shape that annotator happened to use – a literal at a call site, a constant, or nothing at all with the check left to whichever constructor met the resource first – and the refusal a reader got for the wrong resource type differed accordingly.

This is the ANNOTATOR’s copy, not the only one: the wildcard expansion in annotation_config keys the same fact on annotator NAME rather than class (gain#1266), and the web editor states it again per configuration field. What is gone is the five different shapes it took inside the annotators. The wildcard map stays a separate statement on purpose – whether a name expands a wildcard is the annotation layer’s policy, not a property of the annotator (docs/adr/0029-wildcard-expandability-is-parser-policy.md, gain#1334) – and a test pins the two against each other.

The first element is the preferred spelling. A tuple rather than a set for that reason, as FRAGMENT_SCORE_TYPES is one: the order is rendered into the refusal, and AnnotationConfigParser.WILDCARD_RESOURCE_TYPES is pinned against element zero rather than against membership (why, in ADR 0029). So an annotator that comes to accept a further spelling APPENDS it.

Two annotators accept two spellings; each warns from the constructor that opens the resource, which still runs after this check passes the spelling through.

Empty means the annotator does not constrain its resource type – the default, because most annotators (effect_annotator, liftover_annotator, chrom_mapping, …) have no single typed resource to constrain. Those never call resolve_resource(), which refuses an empty declaration rather than rejecting every type in turn.

abstractmethod _do_annotate(annotatable: Annotatable, context: dict[str, Any]) → AnnotatedValues[source]

Annotate the annotatable.

Answers an AnnotatedValues: keyed by attribute NAME, every value finished. The base hands it back as-is (gain#1134); nothing is reduced or renamed on an annotator’s behalf.

An annotator that folds does it in its score’s own read and pairs the answers back with _pair_aggregated(), or – when the values are its own rather than a record stream – through fold_own_values(). One with nothing to fold answers through _from_sources() or _every().

_do_batch_annotate(annotatables: Sequence[Annotatable | None], contexts: list[dict[str, Any]], batch_work_dir: str | None = None) → list[AnnotatedValues][source]

Annotate a batch of annotatables.

One AnnotatedValues per annotatable, in order, on the same contract as _do_annotate().

_empty_result() → AnnotatedValues[source]

None under every attribute name.

The answer for a None annotatable, and what annotators return when a guard fires – a chromosome the resource does not have, a region past the length cutoff.

_every(value: Any) → AnnotatedValues[source]

value under every attribute’s name.

For the annotators with one thing to say – a lifted-over annotatable, a renamed chromosome – however many attributes expose it. The names are read here, at answer time, for the reason AnnotatedValues states.

_from_sources(values: Mapping[str, Any]) → AnnotatedValues[source]

Source-keyed values answered by attribute name, nothing folded.

The rename the base used to do to every result, kept as something an annotator asks for. The non-folding twin of fold_own_values(): for values that are final as they stand – a point read’s one value per score, a tool’s output row, the allele keys an exact match synthesises – where a list is the answer and must not be reduced. An attribute whose source is absent answers None.

_pair_aggregated(values: Sequence[Any], query_count: int, *, resource_id: str, reduced: Callable[[Attribute], bool], otherwise: Callable[[Attribute], Any]) → AnnotatedValues[source]

Pair a score’s reduced values back onto the attributes, by ORDER.

The one statement of how an annotator turns what its score’s folding read answered into an AnnotatedValues. The read’s tuple is parallel to the queries the annotator built over self._attributes in attribute order, so the attributes are walked again here and each one for which reduced holds takes the next value; every other attribute takes otherwise(attr) – the fragment count, the allele keys – whatever that kind answers beside its reductions. The names are read HERE, never cached beside the queries, for the reason AnnotatedValues states.

One value per query, so as many as there are queries: checked rather than assumed, because the pairing is POSITIONAL and a read that answered a different number would otherwise slide every attribute onto its neighbour’s value. A length compare, not a zip(strict=True): the strict zip needs a second list of names to zip against, and building one costs about three times what the whole annotate call costs (measured, gain#1124).

A read that answers one value per attribute and nothing else pairs through _pair_all() instead.

_pair_all(values: Sequence[Any], *, resource_id: str) → AnnotatedValues[source]

Pair one value per attribute back onto the attributes, by ORDER.

The all-reduced case of _pair_aggregated(): for a read that answers exactly one value per attribute, in attribute order, and nothing beside them. Same count check, for the same reason; the pairing itself is one zip over the two lists.

annotate(annotatable: Annotatable | None, context: dict[str, Any]) → dict[str, Any][source]

Answer through _do_annotate; the empty result for None.

Subclasses implement _do_annotate instead of overriding this: the None annotatable is handled here, so _do_annotate never sees one. A batch-only annotator overrides it to raise NotImplementedError.

property attributes: list[Attribute]

The configured attributes, in configuration order.

With no attributes configured, every spec marked is_default stands in, under its source name.

batch_annotate(annotatables: Sequence[Annotatable | None], contexts: list[dict[str, Any]], batch_work_dir: str | None = None) → Sequence[dict[str, Any]][source]

Answer through _do_batch_annotate: one result per annotatable.

Subclasses with a batched backend override _do_batch_annotate instead, whose default loops _do_annotate and handles the None annotatables itself.

get_attribute_defaults(spec: AttributeSpec) → dict[str, Any][source]

Defaults for spec: an aggregator and parameters.

Empty by default. The constructor consults it for every attribute: the aggregator key becomes the aggregator when the configuration names none, and every other key becomes a parameter that the configuration’s own parameters override. Override it when defaults live somewhere other than the spec – a score resource declares its own, for instance.

open() → Annotator[source]

Create work_dir and mark the annotator open; returns self.

Overrides that open resources call this and return self.

static resolve_input_gene_list(pipeline: AnnotationPipeline, info: AnnotatorInfo) → str[source]

Resolve this annotator’s input_gene_list to an attribute name.

The name of an upstream attribute holding the genes the annotator reads, checked to exist in pipeline and to be marked gene_list – the attribute type every gene list the effect annotators produce carries, and the one the web editor offers for this parameter. A bare object attribute is not enough: object is what any structured attribute is stored as.

Beside resolve_resource() for the same reason it is a method here at all: its two callers, the gene-score and gene-set builders, resolve the name to hand to a constructor that has not run yet. The lookup itself is AnnotationPipeline.resolve_attribute_parameter(), shared with the input_annotatable decorator – see there for why.

classmethod resolve_resource(pipeline: AnnotationPipeline, info: AnnotatorInfo) → GenomicResource[source]

Resolve this annotator’s resource_id to a resource it takes.

A classmethod because two of the five callers – the gene-score and gene-set builders – resolve the resource to hand to a constructor that has not run yet; the other three call it as self.resolve_resource(...) from __init__.

Lives on AnnotatorBase rather than on the genomic SCORE base the free function it replaces used to sit beside: two of those callers are not score annotators, and importing the score machinery to reach a type check would be the wrong dependency.

class gain.annotation.annotator_base.AnnotatedValues[source]

The finished answer of a _do_annotate, keyed by ATTRIBUTE NAME.

The seam’s one shape (gain#1130, gain#1134). The keys are attribute names rather than sources because a source exposed twice with two aggregators has two different finished values, which a source-keyed mapping has nowhere to put. Nothing checks the type at run time; it is a contract stated on _do_annotate’s signature and held by the type checker.

Read the names when you ANSWER, never in ``__init__``. The one statement of the rule every annotator building one of these has to follow, kept here rather than in each of them. A pipeline naming one attribute twice renames the later ones – annotation_factory.resolve_repeated_attributes – and it does so AFTER every annotator has been constructed. So an annotator that captured attr.name while building its queries would key its answers by names the pipeline has since moved away from, and the attributes it renamed would come back empty. Whatever a query list caches, it must not cache names; self._attributes is walked again at annotate time and the names read off it then.

gain.annotation.annotator_base.fold_own_values(attributes: Sequence[Attribute], values: Mapping[str, Any]) → AnnotatedValues[source]

Answer an annotator’s OWN values by attribute, each one folded.

For the annotators whose values are their own rather than a score’s record stream – a gene list, a set intersection, one entry per prediction request – so there is no folding read to move the reduction into. Each attribute takes its source’s value, reduced by the aggregator it names (fold()), under the attribute’s NAME.

Only a list is folded. A scalar, a None, an absent source pass through, as does any attribute naming no aggregator: an aggregator says how to reduce MANY values and there is nothing to reduce. That is what the base’s own fold did before gain#1133 retired it.

This is a function rather than a method for the reason gain#1133 exists: the BASE does not aggregate; an annotator that reduces calls this itself before it answers.

class gain.annotation.docker_annotator.DockerAnnotator(pipeline: AnnotationPipeline | None, info: AnnotatorInfo)[source]

Base class for annotators that use docker containers.

open() → Annotator[source]

Create work_dir and mark the annotator open; returns self.

Overrides that open resources call this and return self.