Writing an annotator
Two classes define what an annotator is.
Annotator is the interface the
pipeline drives: open, annotate, close, and the attributes it produces.
AnnotatorBase is the class an
annotator actually extends: it implements the interface once, on top of two
narrower hooks that a subclass fills in. Every annotator GAIn ships, and
every in-tree plugin, is built on AnnotatorBase; the interface class is
documented so you can read the contract the base keeps, not so you can
implement it yourself.
What the base does for you
The base constructor takes the pipeline and an
AnnotatorInfo — the annotator’s
type, id, configured attributes, parameters and resources, as parsed from the
YAML. It checks each configured attribute against
get_attribute_specs(),
resolves its name, aggregator and parameters (consulting
get_attribute_defaults()),
and requires a work_dir parameter — a directory of the annotator’s own,
which the pipeline injects into every annotator’s parameters before the
factory runs and which
open() creates.
Because get_attribute_specs is called from the base constructor, a
subclass that computes its specs from something — a resource’s score list,
say — sets that something before it calls super().__init__.
Two YAML parameters are handled by the pipeline rather than by the
annotator, and a subclass never sees them: input_annotatable, which
substitutes an annotatable produced earlier in the pipeline for the input
row’s own, and an attribute’s value_transform, which rewrites the answer
on its way into the context. Both are applied by wrapping the built
annotator, so the contract below is stated for the annotator alone.
The contract
A subclass implements two members, and may override a handful more.
Always:
get_attribute_specs()— every attribute the annotator can produce, keyed by source (What an annotator declares)._do_annotate()— produce the attributes for one annotatable. It receives the annotatable (neverNone; the base answers the empty result for that before delegating) and the context, and answers anAnnotatedValues: a mapping keyed by attribute name, every value finished. The base hands it back to the pipeline as it is; nothing is reduced or renamed on the annotator’s behalf.
When there is a reason:
_do_batch_annotate()— the same contract for a sequence of annotatables: oneAnnotatedValuesper input, in input order, the empty result for eachNone. The default loops_do_annotateand is correct for every annotator; override it only when the backend has a genuinely batched path — an external tool run once over a file, a model that predicts a batch at a time. A batch-only annotator, one with no per-item path, overridesannotateto raiseNotImplementedErroras well; thedemo_annotatoradapters and the VEP annotators are the shipped examples.open()andclose()— when the annotator holds a resource. Overrides open what they query, call the base and returnself;closemust be safe on an annotator never opened and safe twice.get_attribute_defaults()— when per-attribute defaults live somewhere other than the spec.used_context_attributes— when the annotator reads attributes an earlier annotator produced. Name them here: the pipeline builds its dependency graph from this tuple, and a reannotation reruns the annotator when a named attribute’s producer changes. Reading the context without declaring it works for a plain annotation run, and silently goes stale under reannotation.ACCEPTED_RESOURCE_TYPESandresolve_resource()— when the annotator takes aresource_idnaming a typed genomic resource. Declare the types it accepts, resolve through the classmethod, and the refusal for a resource of the wrong type is phrased the same way as every other annotator’s.
Never: annotate()
and batch_annotate().
They are the pipeline’s side of the contract, and the base already routes
them to the two hooks. The one exception is the batch-only annotate
above.
Shaping the answer
Every AnnotatedValues is keyed by attribute name, and names must be
read when the annotator answers, never cached beside the queries it built —
What an annotator declares says why. The base provides four helpers that read
self.attributes at answer time so a subclass does not have to:
_every()— one value under every attribute’s name. For an annotator with one thing to say however many attributes expose it: a lifted-over annotatable, a renamed chromosome, a yes/no decision._from_sources()— a source-keyed mapping answered by attribute name, nothing folded. For values that are final as they stand — a point read’s one value per score, a tool’s output row. An absent source answersNone.fold_own_values()— the folding twin of_from_sources, for an annotator whose values are its own rather than a score’s record stream: each attribute’s source value is reduced by the aggregator the attribute names, if it names one and the value is a list, and answered under the attribute’s name._pair_aggregated()and_pair_all()— for a score annotator whose score did the folding during the read, and answered one value per query in attribute order. They pair those values back onto the attributes by position, and refuse a read that answered the wrong number of them.
_empty_result() is
_every(None): the answer for a None annotatable, and what an
annotator returns when a guard fires — a chromosome the resource does not
have, a region past a length cutoff.
A minimal annotator
The smallest complete annotator declares one attribute, reads the context,
and answers through _every. This is the shape of the plugin
GAIn Python interface builds in its fifth section, whose full adapter is
downloadable there:
from typing import Any
from gain.annotation.annotatable import Annotatable
from gain.annotation.annotation_pipeline import (
AnnotationPipeline,
Annotator,
AnnotatorInfo,
AttributeSpec,
)
from gain.annotation.annotator_base import AnnotatedValues, AnnotatorBase
class FollowupAnnotator(AnnotatorBase):
"""Flag a variant for follow-up from attributes already computed."""
def get_attribute_specs(self) -> dict[str, AttributeSpec]:
return {
"followup": AttributeSpec(
source="followup",
value_type="str",
description="Whether the variant is selected for follow-up",
),
}
@property
def used_context_attributes(self) -> tuple[str, ...]:
return ("phyloP7way", "clinical_significance")
def _do_annotate(
self, annotatable: Annotatable, context: dict[str, Any],
) -> AnnotatedValues:
conserved = (context.get("phyloP7way") or 0) > 0
pathogenic = context.get("clinical_significance") == "Pathogenic"
return self._every("yes" if conserved and pathogenic else "no")
def build_followup_annotator(
pipeline: AnnotationPipeline, info: AnnotatorInfo,
) -> Annotator:
return FollowupAnnotator(pipeline, info)
The module-level function at the end is the factory: the callable a
pipeline resolves the annotator’s type to, taking the pipeline and the
AnnotatorInfo and returning a built annotator. Annotator plugins says how
it is registered under a name.
Annotators that wrap a tool
An annotator that runs an external program — in a subprocess, or in a container — is a batch annotator by nature: it writes its inputs to a file or a stream, runs the tool once, and reads the answers back. The three in-tree plugin packages are the worked examples, and are cited here rather than paraphrased:
demo_annotator— four annotators, purpose-built as an example. Itsadaptermodule shows the adapter pattern twice over the same toy tool: once through temporary files in the annotator’swork_dir, once over the tool’s stdin and stdout. Its twodemo_annotate_*_adaptermodules show an adapter that hands a GRR resource — gene models, a reference genome — to a tool that reads files, through a caching repository rooted in thework_dir.vep_annotator— two real annotators over Ensembl VEP in a container, on a shared base that builds VEP’s input and parses its output.spliceai_annotator— a real annotator over a model, with a per-item_do_annotateand a batched_do_batch_annotatethat predicts many requests at once; the one in-tree annotator with both paths.
DockerAnnotator is the base the
containerised ones share: it holds a Docker client, prepares the images the
annotator needs in open, and leaves the run of one batch abstract.
API
- class gain.annotation.annotation_pipeline.Annotator(pipeline: AnnotationPipeline | None, info: AnnotatorInfo)[source]
An annotator produces a set of attributes for a given annotatable.
The pipeline drives the lifecycle:
open()before the firstannotate(),close()once at the end. An annotator may assume it is open when asked to annotate, and does not open itself on demand. Implementations extendAnnotatorBase, which handles configuration and theNoneannotatable, rather than this class directly.- abstractmethod annotate(annotatable: Annotatable | None, context: dict[str, Any]) dict[str, Any][source]
Produce this annotator’s attributes for one annotatable.
Returns a mapping from attribute name (not source) to value, with every attribute in
attributespresent.annotatableisNonewhen the input row has none – an unparsable variant, a liftover that found nothing – and the answer is then every attribute set toNone, never an exception.contextholds the attributes of the annotators before this one: read whatused_context_attributesdeclares and do not write to it – the pipeline merges the returned mapping into it. May assumeopen()has run. An annotator that only works in batches raisesNotImplementedErrorhere and overridesbatch_annotate().
- abstract property attributes: list[Attribute]
The attributes this annotator produces, in output order.
Configured attributes: names, sources, aggregators and parameters already resolved against
get_attribute_specs().
- batch_annotate(annotatables: Sequence[Annotatable | None], contexts: list[dict[str, Any]], batch_work_dir: str | None = None) Iterable[dict[str, Any]][source]
Annotate many annotatables: one result per input, in order.
The default calls
annotate()once per pair, lazily, and is correct for every annotator. Override it only when the backend has a genuinely batched path – an external tool run once over a file, say – and keep the same contract: exactly one result per annotatable, in input order; the empty result for aNoneannotatable;contextsread, not written.batch_work_diris a scratch directory the caller may offer,Nonewhen it does not; the default ignores it.
- close() None[source]
Release what
open()acquired and mark the annotator closed.Safe on an annotator never opened, and safe twice; overrides keep it so and call the base. The pipeline calls it once per annotator and logs, rather than propagates, what it raises.
- abstractmethod get_attribute_specs() dict[str, AttributeSpec][source]
Every attribute this annotator can produce, keyed by source.
The catalogue the configuration is checked against: a configured attribute whose source is not a key here is refused. Independent of the configuration and of
open().AnnotatorBasecalls it from its constructor, so it may use only what the subclass set before delegating there.
- get_info() AnnotatorInfo[source]
The annotator info this annotator was built from.
An
AnnotatorInfo: its type, id, configured attributes, parameters and resources.
- open() Annotator[source]
Acquire resources and mark the annotator open; returns
self.The base only flips the flag. Overrides open the resources they query, call the base and return
self. Opening an already-open annotator must be harmless.
- property resources: list[GenomicResource]
The genomic resources this annotator was configured with.
- property used_context_attributes: tuple[str, ...]
Names of upstream attributes this annotator reads from
context.Empty by default. An annotator that reads an attribute another annotator produced – a gene list, say – names it here: the pipeline builds its dependency graph from this tuple, and a reannotation reruns this annotator when a named attribute’s producer changes. Every name must be an attribute of an earlier annotator in the same pipeline.
- class gain.annotation.annotator_base.AnnotatorBase(pipeline: AnnotationPipeline | None, info: AnnotatorInfo)[source]
Base implementation of the Annotator class.
The class every in-tree annotator extends. Its constructor checks the configured attributes against
get_attribute_specs(), resolves each one’s name, aggregator and parameters (consultingget_attribute_defaults()), and requires awork_dirparameter. A subclass implementsget_attribute_specsand_do_annotate(); overrides_do_batch_annotate()when it has a batched path; and overridesget_attribute_defaults(),open()andclose()when it has defaults or resources.annotate()andbatch_annotate()are left alone, except by a batch-only annotator, which makesannotate()refuse.- ACCEPTED_RESOURCE_TYPES: ClassVar[tuple[str, ...]] = ()
The resource types this annotator’s
resource_idmay name.An annotator that consumes a typed genomic resource states them here, and resolves its resource through
resolve_resource(). Before gain#1329 the same fact was written once per annotator in whatever shape that annotator happened to use – a literal at a call site, a constant, or nothing at all with the check left to whichever constructor met the resource first – and the refusal a reader got for the wrong resource type differed accordingly.This is the ANNOTATOR’s copy, not the only one: the wildcard expansion in
annotation_configkeys the same fact on annotator NAME rather than class (gain#1266), and the web editor states it again per configuration field. What is gone is the five different shapes it took inside the annotators. The wildcard map stays a separate statement on purpose – whether a name expands a wildcard is the annotation layer’s policy, not a property of the annotator (docs/adr/0029-wildcard-expandability-is-parser-policy.md, gain#1334) – and a test pins the two against each other.The first element is the preferred spelling. A tuple rather than a set for that reason, as
FRAGMENT_SCORE_TYPESis one: the order is rendered into the refusal, andAnnotationConfigParser.WILDCARD_RESOURCE_TYPESis pinned against element zero rather than against membership (why, in ADR 0029). So an annotator that comes to accept a further spelling APPENDS it.Two annotators accept two spellings; each warns from the constructor that opens the resource, which still runs after this check passes the spelling through.
Empty means the annotator does not constrain its resource type – the default, because most annotators (
effect_annotator,liftover_annotator,chrom_mapping, …) have no single typed resource to constrain. Those never callresolve_resource(), which refuses an empty declaration rather than rejecting every type in turn.
- abstractmethod _do_annotate(annotatable: Annotatable, context: dict[str, Any]) AnnotatedValues[source]
Annotate the annotatable.
Answers an
AnnotatedValues: keyed by attribute NAME, every value finished. The base hands it back as-is (gain#1134); nothing is reduced or renamed on an annotator’s behalf.An annotator that folds does it in its score’s own read and pairs the answers back with
_pair_aggregated(), or – when the values are its own rather than a record stream – throughfold_own_values(). One with nothing to fold answers through_from_sources()or_every().
- _do_batch_annotate(annotatables: Sequence[Annotatable | None], contexts: list[dict[str, Any]], batch_work_dir: str | None = None) list[AnnotatedValues][source]
Annotate a batch of annotatables.
One
AnnotatedValuesper annotatable, in order, on the same contract as_do_annotate().
- _empty_result() AnnotatedValues[source]
Noneunder every attribute name.The answer for a
Noneannotatable, and what annotators return when a guard fires – a chromosome the resource does not have, a region past the length cutoff.
- _every(value: Any) AnnotatedValues[source]
valueunder every attribute’s name.For the annotators with one thing to say – a lifted-over annotatable, a renamed chromosome – however many attributes expose it. The names are read here, at answer time, for the reason
AnnotatedValuesstates.
- _from_sources(values: Mapping[str, Any]) AnnotatedValues[source]
Source-keyed
valuesanswered by attribute name, nothing folded.The rename the base used to do to every result, kept as something an annotator asks for. The non-folding twin of
fold_own_values(): for values that are final as they stand – a point read’s one value per score, a tool’s output row, the allele keys an exact match synthesises – where alistis the answer and must not be reduced. An attribute whose source is absent answersNone.
- _pair_aggregated(values: Sequence[Any], query_count: int, *, resource_id: str, reduced: Callable[[Attribute], bool], otherwise: Callable[[Attribute], Any]) AnnotatedValues[source]
Pair a score’s reduced
valuesback onto the attributes, by ORDER.The one statement of how an annotator turns what its score’s folding read answered into an
AnnotatedValues. The read’s tuple is parallel to the queries the annotator built overself._attributesin attribute order, so the attributes are walked again here and each one for whichreducedholds takes the next value; every other attribute takesotherwise(attr)– the fragment count, the allele keys – whatever that kind answers beside its reductions. The names are read HERE, never cached beside the queries, for the reasonAnnotatedValuesstates.One value per query, so as many as there are queries: checked rather than assumed, because the pairing is POSITIONAL and a read that answered a different number would otherwise slide every attribute onto its neighbour’s value. A length compare, not a
zip(strict=True): the strict zip needs a second list of names to zip against, and building one costs about three times what the whole annotate call costs (measured, gain#1124).A read that answers one value per attribute and nothing else pairs through
_pair_all()instead.
- _pair_all(values: Sequence[Any], *, resource_id: str) AnnotatedValues[source]
Pair one value per attribute back onto the attributes, by ORDER.
The all-reduced case of
_pair_aggregated(): for a read that answers exactly one value per attribute, in attribute order, and nothing beside them. Same count check, for the same reason; the pairing itself is one zip over the two lists.
- annotate(annotatable: Annotatable | None, context: dict[str, Any]) dict[str, Any][source]
Answer through
_do_annotate; the empty result forNone.Subclasses implement
_do_annotateinstead of overriding this: theNoneannotatable is handled here, so_do_annotatenever sees one. A batch-only annotator overrides it to raiseNotImplementedError.
- property attributes: list[Attribute]
The configured attributes, in configuration order.
With no attributes configured, every spec marked
is_defaultstands in, under its source name.
- batch_annotate(annotatables: Sequence[Annotatable | None], contexts: list[dict[str, Any]], batch_work_dir: str | None = None) Sequence[dict[str, Any]][source]
Answer through
_do_batch_annotate: one result per annotatable.Subclasses with a batched backend override
_do_batch_annotateinstead, whose default loops_do_annotateand handles theNoneannotatables itself.
- get_attribute_defaults(spec: AttributeSpec) dict[str, Any][source]
Defaults for
spec: anaggregatorand parameters.Empty by default. The constructor consults it for every attribute: the
aggregatorkey becomes the aggregator when the configuration names none, and every other key becomes a parameter that the configuration’s own parameters override. Override it when defaults live somewhere other than the spec – a score resource declares its own, for instance.
- open() Annotator[source]
Create
work_dirand mark the annotator open; returnsself.Overrides that open resources call this and return
self.
- static resolve_input_gene_list(pipeline: AnnotationPipeline, info: AnnotatorInfo) str[source]
Resolve this annotator’s
input_gene_listto an attribute name.The name of an upstream attribute holding the genes the annotator reads, checked to exist in
pipelineand to be markedgene_list– the attribute type every gene list the effect annotators produce carries, and the one the web editor offers for this parameter. A bareobjectattribute is not enough:objectis what any structured attribute is stored as.Beside
resolve_resource()for the same reason it is a method here at all: its two callers, the gene-score and gene-set builders, resolve the name to hand to a constructor that has not run yet. The lookup itself isAnnotationPipeline.resolve_attribute_parameter(), shared with theinput_annotatabledecorator – see there for why.
- classmethod resolve_resource(pipeline: AnnotationPipeline, info: AnnotatorInfo) GenomicResource[source]
Resolve this annotator’s
resource_idto a resource it takes.A classmethod because two of the five callers – the gene-score and gene-set builders – resolve the resource to hand to a constructor that has not run yet; the other three call it as
self.resolve_resource(...)from__init__.Lives on
AnnotatorBaserather than on the genomic SCORE base the free function it replaces used to sit beside: two of those callers are not score annotators, and importing the score machinery to reach a type check would be the wrong dependency.
- class gain.annotation.annotator_base.AnnotatedValues[source]
The finished answer of a
_do_annotate, keyed by ATTRIBUTE NAME.The seam’s one shape (gain#1130, gain#1134). The keys are attribute names rather than sources because a source exposed twice with two aggregators has two different finished values, which a source-keyed mapping has nowhere to put. Nothing checks the type at run time; it is a contract stated on
_do_annotate’s signature and held by the type checker.Read the names when you ANSWER, never in ``__init__``. The one statement of the rule every annotator building one of these has to follow, kept here rather than in each of them. A pipeline naming one attribute twice renames the later ones –
annotation_factory.resolve_repeated_attributes– and it does so AFTER every annotator has been constructed. So an annotator that capturedattr.namewhile building its queries would key its answers by names the pipeline has since moved away from, and the attributes it renamed would come back empty. Whatever a query list caches, it must not cache names;self._attributesis walked again at annotate time and the names read off it then.
- gain.annotation.annotator_base.fold_own_values(attributes: Sequence[Attribute], values: Mapping[str, Any]) AnnotatedValues[source]
Answer an annotator’s OWN values by attribute, each one folded.
For the annotators whose values are their own rather than a score’s record stream – a gene list, a set intersection, one entry per prediction request – so there is no folding read to move the reduction into. Each attribute takes its source’s value, reduced by the aggregator it names (
fold()), under the attribute’s NAME.Only a
listis folded. A scalar, aNone, an absent source pass through, as does any attribute naming no aggregator: an aggregator says how to reduce MANY values and there is nothing to reduce. That is what the base’s own fold did before gain#1133 retired it.This is a function rather than a method for the reason gain#1133 exists: the BASE does not aggregate; an annotator that reduces calls this itself before it answers.
- class gain.annotation.docker_annotator.DockerAnnotator(pipeline: AnnotationPipeline | None, info: AnnotatorInfo)[source]
Base class for annotators that use docker containers.