Pipelines

An AnnotationPipeline is an ordered list of annotators over one repository. It is what annotate_tabular and annotate_vcf build from the YAML they are given, and what a Python caller builds from the same YAML with load_pipeline_from_yaml().

Loading

Three loaders share one signature and differ only in where the YAML comes from: load_pipeline_from_yaml() takes the text, load_pipeline_from_file() a path, and load_pipeline_from_file_or_resource() either a path or the id of a GRR resource of type annotation_pipeline — the form the command-line tools accept. Each takes the repository the annotators resolve their resources through, and two keyword arguments:

  • work_dir — the directory under which every annotator gets a subdirectory of its own, A<index>_<type>, injected into its parameters as work_dir. Left unset, a temporary directory is minted with a warning; pass one.

  • allow_repeated_attributes — what to do when two annotators emit the same attribute name. Off, the pipeline refuses to build and names the overlap; on, the later attributes are renamed by suffixing their annotator’s id.

Loading does everything but open. For each entry of the YAML, in order, the loader resolves the entry’s type to a factory (Annotator plugins), calls it with the pipeline and the entry’s AnnotatorInfo, wraps the result for input_annotatable and value_transform if the entry uses them, and refuses any parameter the annotator never read. A configuration fault at any of those steps is raised as an AnnotationConfigurationError naming the annotator.

Running

annotate() takes one annotatable (or None) and an optional context, runs every annotator in order, merging each one’s answer into the context before the next runs, and returns the context. So the result carries every attribute of every annotator, internal ones included — it is the writers, not the pipeline, that drop internal attributes from the output.

batch_annotate() does the same over a sequence, one context per input, calling each annotator’s batch_annotate once with the whole sequence. An annotator with a batched backend runs its tool once per call; every other annotator loops. The optional batch_work_dir is a scratch directory the caller may offer to the batched annotators, relative to their own work_dir.

Both open the pipeline on first use if nothing has opened it; both leave it open. The pipeline is a context manager whose exit closes it — and, as with the resource objects, entering it does not open it, so the spelling that does both is with pipeline.open() as pipeline:. close closes every annotator, logging rather than propagating what any of them raises.

Reading a pipeline

Once built, a pipeline answers questions about itself: get_attributes() lists every attribute in pipeline order, get_attribute_info() finds one by name (the first match, so a later annotator that reuses a name is shadowed), and get_annotator_by_attribute_info() finds the annotator that produces it. get_resource_ids() is the set of every resource any annotator uses — what a caching repository would need to fetch to run the pipeline offline. The AnnotatorInfo of each annotator is available through get_info(), and to_dict on it round-trips to the YAML shape it was parsed from.

Reannotation

A ReannotationPipeline is built from a new pipeline and the previous one and runs only what changed: the annotators new to the pipeline; the annotators whose used_context_attributes name an attribute a new annotator now produces; and the annotators whose internal attributes a new annotator reads, since an internal attribute is not in the previous output and has to be computed again. Everything else is copied from the previous output. That is the reason the tuple matters to an implementer: an annotator that reads the context without declaring what it reads is not rerun when its inputs change.

API

gain.annotation.annotation_factory.load_pipeline_from_yaml(raw: str, grr: GenomicResourceRepo, *, allow_repeated_attributes: bool = False, work_dir: Path | None = None) → AnnotationPipeline[source]

Load an annotation pipeline from a YAML-formatted string.

gain.annotation.annotation_factory.load_pipeline_from_file(raw_path: str, grr: GenomicResourceRepo, *, allow_repeated_attributes: bool = False, work_dir: Path | None = None) → AnnotationPipeline[source]

Load an annotation pipeline from a configuration file.

gain.annotation.annotation_factory.load_pipeline_from_file_or_resource(arg: str, grr: GenomicResourceRepo, *, allow_repeated_attributes: bool = False, work_dir: Path | None = None) → AnnotationPipeline[source]

Load a pipeline from a file path or a GRR resource id.

Tries to interpret arg as a filesystem path first; on miss, falls back to looking it up as a GRR resource of type annotation_pipeline.

class gain.annotation.annotation_pipeline.AnnotationPipeline(repository: GenomicResourceRepo)[source]

Provides annotation pipeline abstraction.

add_annotator(annotator: Annotator) → None[source]

Append an annotator; it runs after every annotator already added.

Adding to an open pipeline does not open the annotator: the pipeline opens its annotators only in open().

annotate(annotatable: Annotatable | None, context: dict | None = None) → dict[source]

Apply all annotators to an annotatable.

batch_annotate(annotatables: Sequence[Annotatable | None], contexts: list[dict] | None = None, batch_work_dir: str | None = None) → list[dict][source]

Apply all annotators to a list of annotatables.

close() → None[source]

Close the annotation pipeline.

get_annotator_by_attribute_info(attribute_info: Attribute) → Annotator | None[source]

The annotator producing attribute_info, or None.

Matched by attribute equality, so pass an attribute obtained from this pipeline – get_attribute_info()’s answer, say.

get_attribute_info(attribute_name: str) → Attribute | None[source]

The attribute named attribute_name, or None.

The first match in pipeline order, so a later annotator that reuses a name is shadowed here.

get_attributes() → list[Attribute][source]

Every attribute every annotator produces, in pipeline order.

get_attributes_by_type(attribute_type: str) → list[Attribute][source]

The attributes of one attribute_type, in pipeline order.

Attributes without a spec are skipped.

get_info() → list[AnnotatorInfo][source]

The annotator info of every annotator, in pipeline order.

One AnnotatorInfo per annotator, as each was built from.

get_resource_ids() → set[str][source]

The ids of every resource any annotator uses, as one set.

open() → AnnotationPipeline[source]

Open all annotators in the pipeline and mark it as open.

print() → None[source]

Print the annotation pipeline.

resolve_attribute_parameter(info: AnnotatorInfo, parameter: str, *, expected_attribute_type: str) → str[source]

Resolve info’s parameter to the name of one of my attributes.

The parameter names an upstream attribute the annotator reads. Refused, as a ValueError, when info has no such parameter, when no annotator in the pipeline produces an attribute of that name, or when the attribute’s AttributeSpec.attribute_type is not expected_attribute_type.

One implementation rather than one per caller because the copies it replaces had drifted apart – a typo fixed twice (gain#1170), a misspelling fixed in one copy and left in the other (gain#1280), and the listing of available attributes in the refusal present in one copy and absent from the others (gain#1490).

class gain.annotation.annotation_config.AnnotatorInfo(_type: str, attributes: list[AttributeConfig], parameters: ParamsUsageMonitor | dict[str, Any], documentation: str = '', resources: list[GenomicResource] | None = None, annotator_id: str = 'N/A')[source]

Defines annotator configuration.

to_dict() → dict[str, Any][source]

Convert annotator info to a configuration dictionary.