Pipelines
An AnnotationPipeline is an
ordered list of annotators over one repository. It is what
annotate_tabular and annotate_vcf build from the YAML they are
given, and what a Python caller builds from the same YAML with
load_pipeline_from_yaml().
Loading
Three loaders share one signature and differ only in where the YAML comes
from: load_pipeline_from_yaml()
takes the text, load_pipeline_from_file()
a path, and load_pipeline_from_file_or_resource()
either a path or the id of a GRR resource of type annotation_pipeline —
the form the command-line tools accept. Each takes the repository the
annotators resolve their resources through, and two keyword arguments:
work_dir— the directory under which every annotator gets a subdirectory of its own,A<index>_<type>, injected into its parameters aswork_dir. Left unset, a temporary directory is minted with a warning; pass one.allow_repeated_attributes— what to do when two annotators emit the same attribute name. Off, the pipeline refuses to build and names the overlap; on, the later attributes are renamed by suffixing their annotator’s id.
Loading does everything but open. For each entry of the YAML, in order, the
loader resolves the entry’s type to a factory (Annotator plugins), calls it
with the pipeline and the entry’s AnnotatorInfo,
wraps the result for input_annotatable and value_transform if the
entry uses them, and refuses any parameter the annotator never read. A
configuration fault at any of those steps is raised as an
AnnotationConfigurationError naming the annotator.
Running
annotate()
takes one annotatable (or None) and an optional context, runs every
annotator in order, merging each one’s answer into the context before the
next runs, and returns the context. So the result carries every attribute of
every annotator, internal ones included — it is the writers, not the
pipeline, that drop internal attributes from the output.
batch_annotate()
does the same over a sequence, one context per input, calling each
annotator’s batch_annotate once with the whole sequence. An annotator
with a batched backend runs its tool once per call; every other annotator
loops. The optional batch_work_dir is a scratch directory the caller may
offer to the batched annotators, relative to their own work_dir.
Both open the pipeline on first use if nothing has opened it; both leave it
open. The pipeline is a context manager whose exit closes it — and, as with
the resource objects, entering it does not open it, so the spelling that
does both is with pipeline.open() as pipeline:. close closes every
annotator, logging rather than propagating what any of them raises.
Reading a pipeline
Once built, a pipeline answers questions about itself:
get_attributes()
lists every attribute in pipeline order,
get_attribute_info()
finds one by name (the first match, so a later annotator that reuses a name
is shadowed), and
get_annotator_by_attribute_info()
finds the annotator that produces it.
get_resource_ids()
is the set of every resource any annotator uses — what a caching repository
would need to fetch to run the pipeline offline. The
AnnotatorInfo of each annotator
is available through
get_info(),
and to_dict on it round-trips to the YAML shape it was parsed from.
Reannotation
A ReannotationPipeline is built from a new pipeline and the previous one
and runs only what changed: the annotators new to the pipeline; the
annotators whose used_context_attributes
name an attribute a new annotator now produces; and the annotators whose
internal attributes a new annotator reads, since an internal attribute is
not in the previous output and has to be computed again. Everything else is
copied from the previous output. That is the reason the tuple matters to an
implementer: an annotator that reads the context without declaring what it
reads is not rerun when its inputs change.
API
- gain.annotation.annotation_factory.load_pipeline_from_yaml(raw: str, grr: GenomicResourceRepo, *, allow_repeated_attributes: bool = False, work_dir: Path | None = None) AnnotationPipeline[source]
Load an annotation pipeline from a YAML-formatted string.
- gain.annotation.annotation_factory.load_pipeline_from_file(raw_path: str, grr: GenomicResourceRepo, *, allow_repeated_attributes: bool = False, work_dir: Path | None = None) AnnotationPipeline[source]
Load an annotation pipeline from a configuration file.
- gain.annotation.annotation_factory.load_pipeline_from_file_or_resource(arg: str, grr: GenomicResourceRepo, *, allow_repeated_attributes: bool = False, work_dir: Path | None = None) AnnotationPipeline[source]
Load a pipeline from a file path or a GRR resource id.
Tries to interpret
argas a filesystem path first; on miss, falls back to looking it up as a GRR resource of typeannotation_pipeline.
- class gain.annotation.annotation_pipeline.AnnotationPipeline(repository: GenomicResourceRepo)[source]
Provides annotation pipeline abstraction.
- add_annotator(annotator: Annotator) None[source]
Append an annotator; it runs after every annotator already added.
Adding to an open pipeline does not open the annotator: the pipeline opens its annotators only in
open().
- annotate(annotatable: Annotatable | None, context: dict | None = None) dict[source]
Apply all annotators to an annotatable.
- batch_annotate(annotatables: Sequence[Annotatable | None], contexts: list[dict] | None = None, batch_work_dir: str | None = None) list[dict][source]
Apply all annotators to a list of annotatables.
- get_annotator_by_attribute_info(attribute_info: Attribute) Annotator | None[source]
The annotator producing
attribute_info, orNone.Matched by attribute equality, so pass an attribute obtained from this pipeline –
get_attribute_info()’s answer, say.
- get_attribute_info(attribute_name: str) Attribute | None[source]
The attribute named
attribute_name, orNone.The first match in pipeline order, so a later annotator that reuses a name is shadowed here.
- get_attributes() list[Attribute][source]
Every attribute every annotator produces, in pipeline order.
- get_attributes_by_type(attribute_type: str) list[Attribute][source]
The attributes of one
attribute_type, in pipeline order.Attributes without a spec are skipped.
- get_info() list[AnnotatorInfo][source]
The annotator info of every annotator, in pipeline order.
One
AnnotatorInfoper annotator, as each was built from.
- open() AnnotationPipeline[source]
Open all annotators in the pipeline and mark it as open.
- resolve_attribute_parameter(info: AnnotatorInfo, parameter: str, *, expected_attribute_type: str) str[source]
Resolve
info’sparameterto the name of one of my attributes.The parameter names an upstream attribute the annotator reads. Refused, as a
ValueError, wheninfohas no such parameter, when no annotator in the pipeline produces an attribute of that name, or when the attribute’sAttributeSpec.attribute_typeis notexpected_attribute_type.One implementation rather than one per caller because the copies it replaces had drifted apart – a typo fixed twice (gain#1170), a misspelling fixed in one copy and left in the other (gain#1280), and the listing of available attributes in the refusal present in one copy and absent from the others (gain#1490).
- class gain.annotation.annotation_config.AnnotatorInfo(_type: str, attributes: list[AttributeConfig], parameters: ParamsUsageMonitor | dict[str, Any], documentation: str = '', resources: list[GenomicResource] | None = None, annotator_id: str = 'N/A')[source]
Defines annotator configuration.