MOdular SHotgun metagenome Pipelines with Integrated provenance Tracking: QIIME 2 plugin gor metagenome analysis withtools for genome binning and functional annotation.
- version:
2026.7.0 - website: https://
github .com /bokulich -lab /q2 -annotate - user support:
- Please post to the QIIME 2 forum for help with this plugin: https://
forum .qiime2 .org
Actions¶
| Name | Type | Short Description |
|---|---|---|
| -classify-kraken2 | method | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| collate-kraken2-reports | method | Collate kraken2 reports. |
| collate-kraken2-outputs | method | Collate kraken2 outputs. |
| estimate-bracken | method | Perform read abundance re-estimation using Bracken. |
| build-kraken-db | method | Build Kraken 2 database. |
| inspect-kraken2-db | method | Inspect a Kraken 2 database. |
| kraken2-to-features | method | Select downstream features from Kraken 2. |
| kraken2-to-mag-features | method | Select downstream MAG features from Kraken 2. |
| build-custom-diamond-db | method | Create a DIAMOND formatted reference database from a FASTA input file. |
| fetch-eggnog-db | method | Fetch the databases necessary to run the eggnog-annotate action. |
| fetch-diamond-db | method | Fetch the complete Diamond database necessary to run the eggnog-diamond-search action. |
| fetch-eggnog-proteins | method | Fetch the databases necessary to run the build-eggnog-diamond-db action. |
| fetch-ncbi-taxonomy | method | Fetch NCBI reference taxonomy. |
| build-eggnog-diamond-db | method | Create a DIAMOND formatted reference database for the specified taxon. |
| -eggnog-diamond-search | method | Run eggNOG search using Diamond aligner. |
| -eggnog-hmmer-search | method | Run eggNOG search using HMMER aligner. |
| -eggnog-feature-table | method | Create an eggnog table. |
| -eggnog-annotate | method | Annotate orthologs against eggNOG database. |
| predict-genes-prodigal | method | Predict gene sequences from MAGs or contigs using Prodigal. |
| fetch-kaiju-db | method | Fetch Kaiju database. |
| -classify-kaiju | method | Classify sequences using Kaiju. |
| fetch-eggnog-hmmer-db | method | Fetch the taxon specific database necessary to run the eggnog-hmmer-search action. |
| extract-annotations | method | Extract annotation frequencies from all annotations. |
| -filter-kraken2-reports-by-abundance | method | Filter kraken2 reports by relative abundance. |
| -filter-kraken2-results-by-metadata | method | Filter Kraken2 reports and outputs. |
| -merge-kraken2-results | method | Merge kraken2 reports and outputs. |
| -align-outputs-with-reports | method | Align unfiltered kraken2 outputs with filtered kraken2 reports. |
| -filter-reads-kraken2 | method | Filter Kraken2-classified reads by taxonomy. |
| -visualize-collapsed-contigs | visualizer | Visualize collapsed contig abundances |
| classify-kraken2 | pipeline | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| map-taxonomy-to-contigs | pipeline | Map contig IDs to taxonomy strings from Kraken 2. |
| collapse-contigs | pipeline | Collapse the contig abundances based on taxonomy. |
| search-orthologs-diamond | pipeline | Run eggNOG search using diamond aligner. |
| search-orthologs-hmmer | pipeline | Run eggNOG search using HMMER aligner. |
| map-eggnog | pipeline | Annotate orthologs against eggNOG database. |
| classify-kaiju | pipeline | Classify sequences using Kaiju. |
| filter-kraken2-results | pipeline | Filter kraken2 reports and outputs by metadata and abundance. |
| filter-reads-kraken2 | pipeline | Filter Kraken2-classified reads by taxonomy. |
Artifact Classes¶
EggnogHmmerIdmap |
Formats¶
EggnogHmmerIdmapFileFmt |
EggnogHmmerIdmapDirectoryFmt |
annotate -classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs]|FeatureData[MAG]|SampleData[MAGs] Sequences to be classified. Single-/paired-end reads, contigs, or assembled MAGs can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate collate-kraken2-reports¶
Collates kraken2 reports.
Inputs¶
- reports:
List[SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_reports:
SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate collate-kraken2-outputs¶
Collates kraken2 outputs.
Inputs¶
- outputs:
List[SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_outputs:
SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate estimate-bracken¶
This method uses Bracken to re-estimate read abundances. Only available on Linux platforms.
Citations¶
Inputs¶
- kraken2_reports:
SampleData[Kraken2Report % Properties('reads')] Reports produced by Kraken2.[required]
- db:
BrackenDB Bracken database.[required]
Parameters¶
- threshold:
Int%Range(0, None) Bracken: number of reads required PRIOR to abundance estimation to perform re-estimation.[default:
0]- read_len:
Int%Range(0, None) Bracken: the ideal length of reads in your sample. For paired end data (e.g., 2x150) this should be set to the length of the single-end reads (e.g., 150).[default:
100]- level:
Str%Choices('D', 'P', 'C', 'O', 'F', 'G', 'S') Bracken: specifies the taxonomic rank to analyze. Each classification at this specified rank will receive an estimated number of reads belonging to that rank after abundance estimation.[default:
'S']- include_unclassified:
Bool Bracken does not include the unclassified read counts in the feature table. Set this to True to include those regardless.[default:
True]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('bracken')] Reports modified by Bracken.[required]
- taxonomy:
FeatureData[Taxonomy] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
annotate build-kraken-db¶
This method builds Kraken 2 and Bracken databases either (1) from provided DNA sequences to build a custom database, or (2) simply fetches pre-built versions from an online resource.
Citations¶
Wood et al., 2019; Lu et al., 2017
Inputs¶
- seqs:
List[FeatureData[Sequence]] Sequences to be added to the Kraken 2 database.[optional]
Parameters¶
- collection:
Str%Choices('viral', 'minusb', 'standard', 'standard8', 'standard16', 'pluspf', 'pluspf8', 'pluspf16', 'pluspfp', 'pluspfp8', 'pluspfp16', 'eupathdb', 'nt', 'corent', 'gtdb', 'greengenes', 'rdp', 'silva132', 'silva138') Name of the database collection to be fetched. Please check https://
benlangmead .github .io /aws -indexes /k2 for the description of the available options.[optional] - threads:
Int%Range(1, None) Number of threads. Only applicable when building a custom database.[default:
1]- kmer_len:
Int%Range(1, None) K-mer length in bp/aa.[default:
35]- minimizer_len:
Int%Range(1, None) Minimizer length in bp/aa.[default:
31]- minimizer_spaces:
Int%Range(1, None) Number of characters in minimizer that are ignored in comparisons.[default:
7]- no_masking:
Bool Avoid masking low-complexity sequences prior to building; masking requires dustmasker or segmasker to be installed in PATH.[default:
False]- max_db_size:
Int%Range(0, None) Maximum number of bytes for Kraken 2 hash table; if the estimator determines more would normally be needed, the reference library will be downsampled to fit.[default:
0]- use_ftp:
Bool Use FTP for downloading instead of RSYNC.[default:
False]- load_factor:
Float%Range(0, 1) Proportion of the hash table to be populated.[default:
0.7]- fast_build:
Bool Do not require database to be deterministically built when using multiple threads. This is faster, but does introduce variability in minimizer/LCA pairs.[default:
False]- read_len:
List[Int%Range(1, None)] Ideal read lengths to be used while building the Bracken database.[optional]
Outputs¶
annotate inspect-kraken2-db¶
This method generates a report of identical format to those generated by classify_kraken2, with a slightly different interpretation. Instead of reporting the number of inputs classified to a taxon/clade, the report displays the number of minimizers mapped to each taxon/clade.
Citations¶
Inputs¶
- db:
Kraken2DB The Kraken 2 database for which to generate the report.[required]
Parameters¶
Outputs¶
- report:
Kraken2DBReport The report of the supplied database.[required]
annotate kraken2-to-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
SampleData[Kraken2Report] Per-sample Kraken 2 reports.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- table:
FeatureTable[PresenceAbsence] A presence/absence table of selected features. The features are not of even ranks, but will be the most specific rank available.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate kraken2-to-mag-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
FeatureData[Kraken2Report % Properties('mags')] Per-sample Kraken 2 reports.[required]
- outputs:
FeatureData[Kraken2Output % Properties('mags')] Per-sample Kraken 2 output files.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate build-custom-diamond-db¶
Creates an artifact containing a binary DIAMOND database file (ref_db.dmnd) from a protein reference database file in FASTA format.
Citations¶
Inputs¶
- seqs:
FeatureData[ProteinSequence] Protein reference database.[required]
- taxonomy:
ReferenceDB[NCBITaxonomy] Reference taxonomy, needed to provide taxonomy features.[optional]
Parameters¶
- threads:
Int%Range(1, None) Number of CPU threads.[default:
1]- file_buffer_size:
Int%Range(1, None) File buffer size in bytes.[default:
67108864]- ignore_warnings:
Bool Ignore warnings.[default:
False]- no_parse_seqids:
Bool Print raw seqids without parsing.[default:
False]
Outputs¶
- db:
ReferenceDB[Diamond] DIAMOND database.[required]
annotate fetch-eggnog-db¶
Downloads EggNOG reference database using the download_eggnog_data.py script from eggNOG. Here, this script downloads 3 files and stores them in the output artifact. At least 80 GB of storage space is required to run this action.
Citations¶
Outputs¶
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
annotate fetch-diamond-db¶
Downloads Diamond reference database. This action downloads 1 file (ref_db.dmnd). At least 18 GB of storage space is required to run this action.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database.[required]
annotate fetch-eggnog-proteins¶
Downloads eggNOG proteome database. This script downloads 2 files (e5.proteomes.faa and e5.taxid_info.tsv) and creates and artifact with them. At least 18 GB of storage space is required to run this action.
Citations¶
Outputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information.[required]
annotate fetch-ncbi-taxonomy¶
Downloads NCBI reference taxonomy from the NCBI FTP server. The resulting artifact is required by the build-custom-diamond-db action if one wishes to create a Diamond data base with taxonomy features. At least 30 GB of storage space is required to run this action.
Citations¶
National Center for Biotechnology Information (NCBI), n.d.
Outputs¶
- taxonomy:
ReferenceDB[NCBITaxonomy] NCBI reference taxonomy.[required]
annotate build-eggnog-diamond-db¶
Creates a DIAMOND database which contains the protein sequences that belong to the specified taxon.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information (generated through the
fetch-eggnog-proteinsaction).[required]
Parameters¶
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database for the specified taxon.[required]
annotate -eggnog-diamond-search¶
This method performs the steps by which we find our possible target sequences to annotate using the Diamond search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for ortholog hits.[required]
- db:
ReferenceDB[Diamond] Diamond database.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-hmmer-search¶
This method performs the steps by which we find our possible target sequences to annotate using the HMMER search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-feature-table¶
Create an eggnog table.
Inputs¶
- seed_orthologs:
SampleData[Orthologs] Sequence data to be turned into an eggnog feature table.[required]
Outputs¶
- table:
FeatureTable[Frequency] <no description>[required]
annotate -eggnog-annotate¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] <no description>[required]
- db:
ReferenceDB[Eggnog] <no description>[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggNOG database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] <no description>[required]
annotate predict-genes-prodigal¶
Prodigal (PROkaryotic DYnamic programming Gene-finding ALgorithm), a gene prediction algorithm designed for improved gene structure prediction, translation initiation site recognition, and reduced false positives in bacterial and archaeal genomes.
Citations¶
Inputs¶
- seqs:
FeatureData[MAG]|SampleData[MAGs]|SampleData[Contigs] MAGs or contigs for which one wishes to predict genes.[required]
Parameters¶
- translation_table_number:
Str%Choices('1', '2', '3', '4', '5', '6', '9', '10', '11', '12', '13', '14', '15', '16', '21', '22', '23', '24', '25') Translation table to be used to translate genes into sequences of amino acids. See https://
www .ncbi .nlm .nih .gov /Taxonomy /Utils /wprintgc .cgi for reference.[default: '11']- mode:
Str%Choices('single', 'meta') Gene prediction mode. 'single' is suitable for single genome analysis (e.g., MAGs), 'meta' is suitable for metagenome analysis (e.g., contigs from mixed communities).[default:
'single']- closed:
Bool Treat sequences as complete genomes with closed ends. Use this for finished genomes where the sequences represent complete chromosomes or plasmids.[default:
False]- no_shine_dalgarno:
Bool Bypass Shine-Dalgarno trainer and use a more generic model. Useful for virus, phage, or plasmid sequences that may not follow standard prokaryotic gene patterns.[default:
False]- mask:
Bool Treat runs of N as masked sequence. Useful for assemblies that contain gap regions represented by stretches of N nucleotides.[default:
False]
Outputs¶
- loci:
GenomeData[Loci] Gene coordinates files (one per MAG or sample) listing the location of each predicted gene as well as some additional scoring information.[required]
- genes:
GenomeData[Genes] Fasta files (one per MAG or sample) with the nucleotide sequences of the predicted genes.[required]
- proteins:
GenomeData[Proteins] Fasta files (one per MAG or sample) with the protein translation of the predicted genes.[required]
annotate fetch-kaiju-db¶
This method fetches the latest Kaiju database from Kaiju's web server.
Citations¶
Parameters¶
- database_type:
Str%Choices('nr', 'nr_euk', 'refseq', 'refseq_ref', 'refseq_nr', 'fungi', 'viruses', 'plasmids', 'progenomes', 'rvdb') Type of database to be downloaded. For more information on available types please see the list on Kaiju's web server: https://
bioinformatics -centre .github .io /kaiju /downloads .html[required]
Outputs¶
- db:
KaijuDB Kaiju database.[required]
annotate -classify-kaiju¶
This method uses Kaiju to perform taxonomic classification of DNA sequence reads or contigs.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate fetch-eggnog-hmmer-db¶
Downloads Profile HMM database for the specified taxon.
Citations¶
Huerta-Cepas et al., 2019; HMMER, 2024
Parameters¶
Outputs¶
- idmap:
EggnogHmmerIdmap%Properties('eggnog') List of protein families in
hmm_db.[required]- hmm_db:
ProfileHMM[MultipleProtein]%Properties('eggnog') Collection of Profile HMMs.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein]%Properties('eggnog') Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins]%Properties('eggnog') Seed alignments for the protein families in
hmm_db.[required]
annotate extract-annotations¶
This method extract a specific annotation from the table generated by EggNOG and calculates its frequencies across all contigs (annotation_counts_per_contig) and MAGs (annotation_counts_per_genome).
Inputs¶
- ortholog_annotations:
GenomeData[NOG] Ortholog annotations.[required]
Parameters¶
- annotation:
Str%Choices('cog', 'caz', 'kegg_ko', 'kegg_pathway', 'kegg_reaction', 'kegg_module', 'brite', 'ec') Annotation to extract.[required]
- max_evalue:
Float%Range(0, None) <no description>[default:
1.0]- min_score:
Float%Range(0, None) <no description>[default:
0.0]
Outputs¶
- annotation_counts_per_genome:
FeatureTable[Frequency] Feature table with frequency of each annotation per genome.[required]
- annotation_map:
FeatureMap[FunctionToContigs] Feature map with function to contigs mapping.[required]
- annotation_counts_per_contig:
FeatureTable[Frequency] Feature table with frequency of each annotation per contig.[required]
annotate -filter-kraken2-reports-by-abundance¶
Filters kraken2 reports on a per-taxon basis by relative abundance (relative frequency). Useful for removing suspected spurious classifications.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter by relative abundance.[required]
Parameters¶
- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report.[required]
- remove_empty:
Bool If True, reports with only unclassified reads remaining will be removed from the filtered data.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The relative abundance-filtered kraken2 reports[required]
annotate -filter-kraken2-results-by-metadata¶
Filter Kraken2 reports and outputs based on metadata or remove reports with 100% unclassified reads.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The Kraken reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The Kraken outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] <no description>[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] <no description>[required]
annotate -merge-kraken2-results¶
Merge multiple kraken2 reports and outputs such that the results contain a union of the samples represented in the inputs. If sample IDs overlap across the inputs, these reports and outputs will be processed into a single report or output per sample ID.
Inputs¶
- reports:
List[SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')]] The kraken2 reports to merge. Only reports with the same sample ID are merged into one report.[required]
- outputs:
List[SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')]] The kraken2 outputs to merge. Only outputs with the same sample ID are merged into one output.[required]
Outputs¶
- merged_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The merged kraken2 reports.[required]
- merged_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The merged kraken2 outputs.[required]
annotate -align-outputs-with-reports¶
Inputs¶
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to align with the filtered reports.[required]
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
Outputs¶
- aligned_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The report-aligned filtered kraken2 outputs.[required]
annotate -filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
annotate -visualize-collapsed-contigs¶
Generates a visualization showing histograms of contig abundance distributions per taxon per sample from the original (pre-collapsed) table.
Inputs¶
- table:
FeatureTable[Frequency] Original feature table with contig IDs as feature IDs (before collapsing).[required]
- collapsed_table:
FeatureTable[Frequency] Feature table with contig IDs collapsed to taxonomy IDs.[required]
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings instead of IDs.[optional]
Outputs¶
- visualization:
Visualization <no description>[required]
annotate classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
List[SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]]|List[SampleData[Contigs]]|List[FeatureData[MAG]]|List[SampleData[MAGs]] Sequences to be classified. Single-/paired-end reads,contigs, or assembled MAGs, can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate map-taxonomy-to-contigs¶
Maps contig IDs to their full taxonomy strings based on Kraken 2 classifications. This action processes Kraken 2 reports and outputs produced from contig sequences to create a taxonomy mapping where each contig ID is associated with its assigned taxonomy.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('contigs')] Per-sample Kraken 2 reports for contigs.[required]
- outputs:
SampleData[Kraken2Output % Properties('contigs')] Per-sample Kraken 2 output files for contigs.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- feature_map:
FeatureMap[TaxonomyToContigs] Taxonomy assignments for contigs. Each contig ID is mapped to its full taxonomy string based on Kraken2 classifications. Unclassified contigs are assigned 'd__Unclassified'. Assignments below the provided coverage threshold will be treated as 'd__Unclassified'.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate collapse-contigs¶
This action collapses contig abundances based on their taxonomy assignments. Contig abundances for contigs with the same taxonomy assignment will be averaged.
Inputs¶
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- table:
FeatureTable[Frequency] Table of contig abundances per sample. Feature IDs will be replaced with taxonomy strings.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings in the visualization instead of taxonomy IDs.[optional]
Outputs¶
- collapsed_table:
FeatureTable[Frequency] Contig abundance table with contig IDs replaced by their taxonomy strings. Contigs with the same taxonomy assignment will have their abundances averaged.[required]
- visualization:
Visualization Interactive visualization showing histograms of contig abundance distributions per taxon per sample.[required]
annotate search-orthologs-diamond¶
Use Diamond and eggNOG to align contig or MAG sequences against the Diamond database.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits using the Diamond Database[required]
- db:
ReferenceDB[Diamond] The filepath to an artifact containing the Diamond database[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate search-orthologs-hmmer¶
This method uses HMMER to find possible target sequences to annotate with eggNOG-mapper.
Citations¶
HMMER, 2024; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of profile HMMs in binary format and indexed.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate map-eggnog¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] BLAST6-like table(s) describing the identified orthologs.[required]
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggnog database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] Annotated hits.[required]
annotate classify-kaiju¶
This method uses Kaiju to perform taxonomic classification.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate filter-kraken2-results¶
Filter kraken2 reports and outputs by sample metadata, and/or filter classified taxa by relative abundance.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report. If a taxon is filtered from the report, its associated read counts are removed entirely from the report (i.e., the subtraction of those counts is propagated to parent taxonomic groupings).[optional]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The filtered kraken2 outputs.[required]
annotate filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[default:
1]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
MOdular SHotgun metagenome Pipelines with Integrated provenance Tracking: QIIME 2 plugin gor metagenome analysis withtools for genome binning and functional annotation.
- version:
2026.7.0 - website: https://
github .com /bokulich -lab /q2 -annotate - user support:
- Please post to the QIIME 2 forum for help with this plugin: https://
forum .qiime2 .org
Actions¶
| Name | Type | Short Description |
|---|---|---|
| -classify-kraken2 | method | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| collate-kraken2-reports | method | Collate kraken2 reports. |
| collate-kraken2-outputs | method | Collate kraken2 outputs. |
| estimate-bracken | method | Perform read abundance re-estimation using Bracken. |
| build-kraken-db | method | Build Kraken 2 database. |
| inspect-kraken2-db | method | Inspect a Kraken 2 database. |
| kraken2-to-features | method | Select downstream features from Kraken 2. |
| kraken2-to-mag-features | method | Select downstream MAG features from Kraken 2. |
| build-custom-diamond-db | method | Create a DIAMOND formatted reference database from a FASTA input file. |
| fetch-eggnog-db | method | Fetch the databases necessary to run the eggnog-annotate action. |
| fetch-diamond-db | method | Fetch the complete Diamond database necessary to run the eggnog-diamond-search action. |
| fetch-eggnog-proteins | method | Fetch the databases necessary to run the build-eggnog-diamond-db action. |
| fetch-ncbi-taxonomy | method | Fetch NCBI reference taxonomy. |
| build-eggnog-diamond-db | method | Create a DIAMOND formatted reference database for the specified taxon. |
| -eggnog-diamond-search | method | Run eggNOG search using Diamond aligner. |
| -eggnog-hmmer-search | method | Run eggNOG search using HMMER aligner. |
| -eggnog-feature-table | method | Create an eggnog table. |
| -eggnog-annotate | method | Annotate orthologs against eggNOG database. |
| predict-genes-prodigal | method | Predict gene sequences from MAGs or contigs using Prodigal. |
| fetch-kaiju-db | method | Fetch Kaiju database. |
| -classify-kaiju | method | Classify sequences using Kaiju. |
| fetch-eggnog-hmmer-db | method | Fetch the taxon specific database necessary to run the eggnog-hmmer-search action. |
| extract-annotations | method | Extract annotation frequencies from all annotations. |
| -filter-kraken2-reports-by-abundance | method | Filter kraken2 reports by relative abundance. |
| -filter-kraken2-results-by-metadata | method | Filter Kraken2 reports and outputs. |
| -merge-kraken2-results | method | Merge kraken2 reports and outputs. |
| -align-outputs-with-reports | method | Align unfiltered kraken2 outputs with filtered kraken2 reports. |
| -filter-reads-kraken2 | method | Filter Kraken2-classified reads by taxonomy. |
| -visualize-collapsed-contigs | visualizer | Visualize collapsed contig abundances |
| classify-kraken2 | pipeline | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| map-taxonomy-to-contigs | pipeline | Map contig IDs to taxonomy strings from Kraken 2. |
| collapse-contigs | pipeline | Collapse the contig abundances based on taxonomy. |
| search-orthologs-diamond | pipeline | Run eggNOG search using diamond aligner. |
| search-orthologs-hmmer | pipeline | Run eggNOG search using HMMER aligner. |
| map-eggnog | pipeline | Annotate orthologs against eggNOG database. |
| classify-kaiju | pipeline | Classify sequences using Kaiju. |
| filter-kraken2-results | pipeline | Filter kraken2 reports and outputs by metadata and abundance. |
| filter-reads-kraken2 | pipeline | Filter Kraken2-classified reads by taxonomy. |
Artifact Classes¶
EggnogHmmerIdmap |
Formats¶
EggnogHmmerIdmapFileFmt |
EggnogHmmerIdmapDirectoryFmt |
annotate -classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs]|FeatureData[MAG]|SampleData[MAGs] Sequences to be classified. Single-/paired-end reads, contigs, or assembled MAGs can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate collate-kraken2-reports¶
Collates kraken2 reports.
Inputs¶
- reports:
List[SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_reports:
SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate collate-kraken2-outputs¶
Collates kraken2 outputs.
Inputs¶
- outputs:
List[SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_outputs:
SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate estimate-bracken¶
This method uses Bracken to re-estimate read abundances. Only available on Linux platforms.
Citations¶
Inputs¶
- kraken2_reports:
SampleData[Kraken2Report % Properties('reads')] Reports produced by Kraken2.[required]
- db:
BrackenDB Bracken database.[required]
Parameters¶
- threshold:
Int%Range(0, None) Bracken: number of reads required PRIOR to abundance estimation to perform re-estimation.[default:
0]- read_len:
Int%Range(0, None) Bracken: the ideal length of reads in your sample. For paired end data (e.g., 2x150) this should be set to the length of the single-end reads (e.g., 150).[default:
100]- level:
Str%Choices('D', 'P', 'C', 'O', 'F', 'G', 'S') Bracken: specifies the taxonomic rank to analyze. Each classification at this specified rank will receive an estimated number of reads belonging to that rank after abundance estimation.[default:
'S']- include_unclassified:
Bool Bracken does not include the unclassified read counts in the feature table. Set this to True to include those regardless.[default:
True]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('bracken')] Reports modified by Bracken.[required]
- taxonomy:
FeatureData[Taxonomy] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
annotate build-kraken-db¶
This method builds Kraken 2 and Bracken databases either (1) from provided DNA sequences to build a custom database, or (2) simply fetches pre-built versions from an online resource.
Citations¶
Wood et al., 2019; Lu et al., 2017
Inputs¶
- seqs:
List[FeatureData[Sequence]] Sequences to be added to the Kraken 2 database.[optional]
Parameters¶
- collection:
Str%Choices('viral', 'minusb', 'standard', 'standard8', 'standard16', 'pluspf', 'pluspf8', 'pluspf16', 'pluspfp', 'pluspfp8', 'pluspfp16', 'eupathdb', 'nt', 'corent', 'gtdb', 'greengenes', 'rdp', 'silva132', 'silva138') Name of the database collection to be fetched. Please check https://
benlangmead .github .io /aws -indexes /k2 for the description of the available options.[optional] - threads:
Int%Range(1, None) Number of threads. Only applicable when building a custom database.[default:
1]- kmer_len:
Int%Range(1, None) K-mer length in bp/aa.[default:
35]- minimizer_len:
Int%Range(1, None) Minimizer length in bp/aa.[default:
31]- minimizer_spaces:
Int%Range(1, None) Number of characters in minimizer that are ignored in comparisons.[default:
7]- no_masking:
Bool Avoid masking low-complexity sequences prior to building; masking requires dustmasker or segmasker to be installed in PATH.[default:
False]- max_db_size:
Int%Range(0, None) Maximum number of bytes for Kraken 2 hash table; if the estimator determines more would normally be needed, the reference library will be downsampled to fit.[default:
0]- use_ftp:
Bool Use FTP for downloading instead of RSYNC.[default:
False]- load_factor:
Float%Range(0, 1) Proportion of the hash table to be populated.[default:
0.7]- fast_build:
Bool Do not require database to be deterministically built when using multiple threads. This is faster, but does introduce variability in minimizer/LCA pairs.[default:
False]- read_len:
List[Int%Range(1, None)] Ideal read lengths to be used while building the Bracken database.[optional]
Outputs¶
annotate inspect-kraken2-db¶
This method generates a report of identical format to those generated by classify_kraken2, with a slightly different interpretation. Instead of reporting the number of inputs classified to a taxon/clade, the report displays the number of minimizers mapped to each taxon/clade.
Citations¶
Inputs¶
- db:
Kraken2DB The Kraken 2 database for which to generate the report.[required]
Parameters¶
Outputs¶
- report:
Kraken2DBReport The report of the supplied database.[required]
annotate kraken2-to-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
SampleData[Kraken2Report] Per-sample Kraken 2 reports.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- table:
FeatureTable[PresenceAbsence] A presence/absence table of selected features. The features are not of even ranks, but will be the most specific rank available.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate kraken2-to-mag-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
FeatureData[Kraken2Report % Properties('mags')] Per-sample Kraken 2 reports.[required]
- outputs:
FeatureData[Kraken2Output % Properties('mags')] Per-sample Kraken 2 output files.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate build-custom-diamond-db¶
Creates an artifact containing a binary DIAMOND database file (ref_db.dmnd) from a protein reference database file in FASTA format.
Citations¶
Inputs¶
- seqs:
FeatureData[ProteinSequence] Protein reference database.[required]
- taxonomy:
ReferenceDB[NCBITaxonomy] Reference taxonomy, needed to provide taxonomy features.[optional]
Parameters¶
- threads:
Int%Range(1, None) Number of CPU threads.[default:
1]- file_buffer_size:
Int%Range(1, None) File buffer size in bytes.[default:
67108864]- ignore_warnings:
Bool Ignore warnings.[default:
False]- no_parse_seqids:
Bool Print raw seqids without parsing.[default:
False]
Outputs¶
- db:
ReferenceDB[Diamond] DIAMOND database.[required]
annotate fetch-eggnog-db¶
Downloads EggNOG reference database using the download_eggnog_data.py script from eggNOG. Here, this script downloads 3 files and stores them in the output artifact. At least 80 GB of storage space is required to run this action.
Citations¶
Outputs¶
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
annotate fetch-diamond-db¶
Downloads Diamond reference database. This action downloads 1 file (ref_db.dmnd). At least 18 GB of storage space is required to run this action.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database.[required]
annotate fetch-eggnog-proteins¶
Downloads eggNOG proteome database. This script downloads 2 files (e5.proteomes.faa and e5.taxid_info.tsv) and creates and artifact with them. At least 18 GB of storage space is required to run this action.
Citations¶
Outputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information.[required]
annotate fetch-ncbi-taxonomy¶
Downloads NCBI reference taxonomy from the NCBI FTP server. The resulting artifact is required by the build-custom-diamond-db action if one wishes to create a Diamond data base with taxonomy features. At least 30 GB of storage space is required to run this action.
Citations¶
National Center for Biotechnology Information (NCBI), n.d.
Outputs¶
- taxonomy:
ReferenceDB[NCBITaxonomy] NCBI reference taxonomy.[required]
annotate build-eggnog-diamond-db¶
Creates a DIAMOND database which contains the protein sequences that belong to the specified taxon.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information (generated through the
fetch-eggnog-proteinsaction).[required]
Parameters¶
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database for the specified taxon.[required]
annotate -eggnog-diamond-search¶
This method performs the steps by which we find our possible target sequences to annotate using the Diamond search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for ortholog hits.[required]
- db:
ReferenceDB[Diamond] Diamond database.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-hmmer-search¶
This method performs the steps by which we find our possible target sequences to annotate using the HMMER search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-feature-table¶
Create an eggnog table.
Inputs¶
- seed_orthologs:
SampleData[Orthologs] Sequence data to be turned into an eggnog feature table.[required]
Outputs¶
- table:
FeatureTable[Frequency] <no description>[required]
annotate -eggnog-annotate¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] <no description>[required]
- db:
ReferenceDB[Eggnog] <no description>[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggNOG database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] <no description>[required]
annotate predict-genes-prodigal¶
Prodigal (PROkaryotic DYnamic programming Gene-finding ALgorithm), a gene prediction algorithm designed for improved gene structure prediction, translation initiation site recognition, and reduced false positives in bacterial and archaeal genomes.
Citations¶
Inputs¶
- seqs:
FeatureData[MAG]|SampleData[MAGs]|SampleData[Contigs] MAGs or contigs for which one wishes to predict genes.[required]
Parameters¶
- translation_table_number:
Str%Choices('1', '2', '3', '4', '5', '6', '9', '10', '11', '12', '13', '14', '15', '16', '21', '22', '23', '24', '25') Translation table to be used to translate genes into sequences of amino acids. See https://
www .ncbi .nlm .nih .gov /Taxonomy /Utils /wprintgc .cgi for reference.[default: '11']- mode:
Str%Choices('single', 'meta') Gene prediction mode. 'single' is suitable for single genome analysis (e.g., MAGs), 'meta' is suitable for metagenome analysis (e.g., contigs from mixed communities).[default:
'single']- closed:
Bool Treat sequences as complete genomes with closed ends. Use this for finished genomes where the sequences represent complete chromosomes or plasmids.[default:
False]- no_shine_dalgarno:
Bool Bypass Shine-Dalgarno trainer and use a more generic model. Useful for virus, phage, or plasmid sequences that may not follow standard prokaryotic gene patterns.[default:
False]- mask:
Bool Treat runs of N as masked sequence. Useful for assemblies that contain gap regions represented by stretches of N nucleotides.[default:
False]
Outputs¶
- loci:
GenomeData[Loci] Gene coordinates files (one per MAG or sample) listing the location of each predicted gene as well as some additional scoring information.[required]
- genes:
GenomeData[Genes] Fasta files (one per MAG or sample) with the nucleotide sequences of the predicted genes.[required]
- proteins:
GenomeData[Proteins] Fasta files (one per MAG or sample) with the protein translation of the predicted genes.[required]
annotate fetch-kaiju-db¶
This method fetches the latest Kaiju database from Kaiju's web server.
Citations¶
Parameters¶
- database_type:
Str%Choices('nr', 'nr_euk', 'refseq', 'refseq_ref', 'refseq_nr', 'fungi', 'viruses', 'plasmids', 'progenomes', 'rvdb') Type of database to be downloaded. For more information on available types please see the list on Kaiju's web server: https://
bioinformatics -centre .github .io /kaiju /downloads .html[required]
Outputs¶
- db:
KaijuDB Kaiju database.[required]
annotate -classify-kaiju¶
This method uses Kaiju to perform taxonomic classification of DNA sequence reads or contigs.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate fetch-eggnog-hmmer-db¶
Downloads Profile HMM database for the specified taxon.
Citations¶
Huerta-Cepas et al., 2019; HMMER, 2024
Parameters¶
Outputs¶
- idmap:
EggnogHmmerIdmap%Properties('eggnog') List of protein families in
hmm_db.[required]- hmm_db:
ProfileHMM[MultipleProtein]%Properties('eggnog') Collection of Profile HMMs.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein]%Properties('eggnog') Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins]%Properties('eggnog') Seed alignments for the protein families in
hmm_db.[required]
annotate extract-annotations¶
This method extract a specific annotation from the table generated by EggNOG and calculates its frequencies across all contigs (annotation_counts_per_contig) and MAGs (annotation_counts_per_genome).
Inputs¶
- ortholog_annotations:
GenomeData[NOG] Ortholog annotations.[required]
Parameters¶
- annotation:
Str%Choices('cog', 'caz', 'kegg_ko', 'kegg_pathway', 'kegg_reaction', 'kegg_module', 'brite', 'ec') Annotation to extract.[required]
- max_evalue:
Float%Range(0, None) <no description>[default:
1.0]- min_score:
Float%Range(0, None) <no description>[default:
0.0]
Outputs¶
- annotation_counts_per_genome:
FeatureTable[Frequency] Feature table with frequency of each annotation per genome.[required]
- annotation_map:
FeatureMap[FunctionToContigs] Feature map with function to contigs mapping.[required]
- annotation_counts_per_contig:
FeatureTable[Frequency] Feature table with frequency of each annotation per contig.[required]
annotate -filter-kraken2-reports-by-abundance¶
Filters kraken2 reports on a per-taxon basis by relative abundance (relative frequency). Useful for removing suspected spurious classifications.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter by relative abundance.[required]
Parameters¶
- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report.[required]
- remove_empty:
Bool If True, reports with only unclassified reads remaining will be removed from the filtered data.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The relative abundance-filtered kraken2 reports[required]
annotate -filter-kraken2-results-by-metadata¶
Filter Kraken2 reports and outputs based on metadata or remove reports with 100% unclassified reads.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The Kraken reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The Kraken outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] <no description>[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] <no description>[required]
annotate -merge-kraken2-results¶
Merge multiple kraken2 reports and outputs such that the results contain a union of the samples represented in the inputs. If sample IDs overlap across the inputs, these reports and outputs will be processed into a single report or output per sample ID.
Inputs¶
- reports:
List[SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')]] The kraken2 reports to merge. Only reports with the same sample ID are merged into one report.[required]
- outputs:
List[SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')]] The kraken2 outputs to merge. Only outputs with the same sample ID are merged into one output.[required]
Outputs¶
- merged_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The merged kraken2 reports.[required]
- merged_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The merged kraken2 outputs.[required]
annotate -align-outputs-with-reports¶
Inputs¶
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to align with the filtered reports.[required]
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
Outputs¶
- aligned_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The report-aligned filtered kraken2 outputs.[required]
annotate -filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
annotate -visualize-collapsed-contigs¶
Generates a visualization showing histograms of contig abundance distributions per taxon per sample from the original (pre-collapsed) table.
Inputs¶
- table:
FeatureTable[Frequency] Original feature table with contig IDs as feature IDs (before collapsing).[required]
- collapsed_table:
FeatureTable[Frequency] Feature table with contig IDs collapsed to taxonomy IDs.[required]
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings instead of IDs.[optional]
Outputs¶
- visualization:
Visualization <no description>[required]
annotate classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
List[SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]]|List[SampleData[Contigs]]|List[FeatureData[MAG]]|List[SampleData[MAGs]] Sequences to be classified. Single-/paired-end reads,contigs, or assembled MAGs, can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate map-taxonomy-to-contigs¶
Maps contig IDs to their full taxonomy strings based on Kraken 2 classifications. This action processes Kraken 2 reports and outputs produced from contig sequences to create a taxonomy mapping where each contig ID is associated with its assigned taxonomy.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('contigs')] Per-sample Kraken 2 reports for contigs.[required]
- outputs:
SampleData[Kraken2Output % Properties('contigs')] Per-sample Kraken 2 output files for contigs.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- feature_map:
FeatureMap[TaxonomyToContigs] Taxonomy assignments for contigs. Each contig ID is mapped to its full taxonomy string based on Kraken2 classifications. Unclassified contigs are assigned 'd__Unclassified'. Assignments below the provided coverage threshold will be treated as 'd__Unclassified'.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate collapse-contigs¶
This action collapses contig abundances based on their taxonomy assignments. Contig abundances for contigs with the same taxonomy assignment will be averaged.
Inputs¶
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- table:
FeatureTable[Frequency] Table of contig abundances per sample. Feature IDs will be replaced with taxonomy strings.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings in the visualization instead of taxonomy IDs.[optional]
Outputs¶
- collapsed_table:
FeatureTable[Frequency] Contig abundance table with contig IDs replaced by their taxonomy strings. Contigs with the same taxonomy assignment will have their abundances averaged.[required]
- visualization:
Visualization Interactive visualization showing histograms of contig abundance distributions per taxon per sample.[required]
annotate search-orthologs-diamond¶
Use Diamond and eggNOG to align contig or MAG sequences against the Diamond database.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits using the Diamond Database[required]
- db:
ReferenceDB[Diamond] The filepath to an artifact containing the Diamond database[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate search-orthologs-hmmer¶
This method uses HMMER to find possible target sequences to annotate with eggNOG-mapper.
Citations¶
HMMER, 2024; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of profile HMMs in binary format and indexed.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate map-eggnog¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] BLAST6-like table(s) describing the identified orthologs.[required]
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggnog database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] Annotated hits.[required]
annotate classify-kaiju¶
This method uses Kaiju to perform taxonomic classification.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate filter-kraken2-results¶
Filter kraken2 reports and outputs by sample metadata, and/or filter classified taxa by relative abundance.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report. If a taxon is filtered from the report, its associated read counts are removed entirely from the report (i.e., the subtraction of those counts is propagated to parent taxonomic groupings).[optional]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The filtered kraken2 outputs.[required]
annotate filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[default:
1]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
MOdular SHotgun metagenome Pipelines with Integrated provenance Tracking: QIIME 2 plugin gor metagenome analysis withtools for genome binning and functional annotation.
- version:
2026.7.0 - website: https://
github .com /bokulich -lab /q2 -annotate - user support:
- Please post to the QIIME 2 forum for help with this plugin: https://
forum .qiime2 .org
Actions¶
| Name | Type | Short Description |
|---|---|---|
| -classify-kraken2 | method | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| collate-kraken2-reports | method | Collate kraken2 reports. |
| collate-kraken2-outputs | method | Collate kraken2 outputs. |
| estimate-bracken | method | Perform read abundance re-estimation using Bracken. |
| build-kraken-db | method | Build Kraken 2 database. |
| inspect-kraken2-db | method | Inspect a Kraken 2 database. |
| kraken2-to-features | method | Select downstream features from Kraken 2. |
| kraken2-to-mag-features | method | Select downstream MAG features from Kraken 2. |
| build-custom-diamond-db | method | Create a DIAMOND formatted reference database from a FASTA input file. |
| fetch-eggnog-db | method | Fetch the databases necessary to run the eggnog-annotate action. |
| fetch-diamond-db | method | Fetch the complete Diamond database necessary to run the eggnog-diamond-search action. |
| fetch-eggnog-proteins | method | Fetch the databases necessary to run the build-eggnog-diamond-db action. |
| fetch-ncbi-taxonomy | method | Fetch NCBI reference taxonomy. |
| build-eggnog-diamond-db | method | Create a DIAMOND formatted reference database for the specified taxon. |
| -eggnog-diamond-search | method | Run eggNOG search using Diamond aligner. |
| -eggnog-hmmer-search | method | Run eggNOG search using HMMER aligner. |
| -eggnog-feature-table | method | Create an eggnog table. |
| -eggnog-annotate | method | Annotate orthologs against eggNOG database. |
| predict-genes-prodigal | method | Predict gene sequences from MAGs or contigs using Prodigal. |
| fetch-kaiju-db | method | Fetch Kaiju database. |
| -classify-kaiju | method | Classify sequences using Kaiju. |
| fetch-eggnog-hmmer-db | method | Fetch the taxon specific database necessary to run the eggnog-hmmer-search action. |
| extract-annotations | method | Extract annotation frequencies from all annotations. |
| -filter-kraken2-reports-by-abundance | method | Filter kraken2 reports by relative abundance. |
| -filter-kraken2-results-by-metadata | method | Filter Kraken2 reports and outputs. |
| -merge-kraken2-results | method | Merge kraken2 reports and outputs. |
| -align-outputs-with-reports | method | Align unfiltered kraken2 outputs with filtered kraken2 reports. |
| -filter-reads-kraken2 | method | Filter Kraken2-classified reads by taxonomy. |
| -visualize-collapsed-contigs | visualizer | Visualize collapsed contig abundances |
| classify-kraken2 | pipeline | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| map-taxonomy-to-contigs | pipeline | Map contig IDs to taxonomy strings from Kraken 2. |
| collapse-contigs | pipeline | Collapse the contig abundances based on taxonomy. |
| search-orthologs-diamond | pipeline | Run eggNOG search using diamond aligner. |
| search-orthologs-hmmer | pipeline | Run eggNOG search using HMMER aligner. |
| map-eggnog | pipeline | Annotate orthologs against eggNOG database. |
| classify-kaiju | pipeline | Classify sequences using Kaiju. |
| filter-kraken2-results | pipeline | Filter kraken2 reports and outputs by metadata and abundance. |
| filter-reads-kraken2 | pipeline | Filter Kraken2-classified reads by taxonomy. |
Artifact Classes¶
EggnogHmmerIdmap |
Formats¶
EggnogHmmerIdmapFileFmt |
EggnogHmmerIdmapDirectoryFmt |
annotate -classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs]|FeatureData[MAG]|SampleData[MAGs] Sequences to be classified. Single-/paired-end reads, contigs, or assembled MAGs can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate collate-kraken2-reports¶
Collates kraken2 reports.
Inputs¶
- reports:
List[SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_reports:
SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate collate-kraken2-outputs¶
Collates kraken2 outputs.
Inputs¶
- outputs:
List[SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_outputs:
SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate estimate-bracken¶
This method uses Bracken to re-estimate read abundances. Only available on Linux platforms.
Citations¶
Inputs¶
- kraken2_reports:
SampleData[Kraken2Report % Properties('reads')] Reports produced by Kraken2.[required]
- db:
BrackenDB Bracken database.[required]
Parameters¶
- threshold:
Int%Range(0, None) Bracken: number of reads required PRIOR to abundance estimation to perform re-estimation.[default:
0]- read_len:
Int%Range(0, None) Bracken: the ideal length of reads in your sample. For paired end data (e.g., 2x150) this should be set to the length of the single-end reads (e.g., 150).[default:
100]- level:
Str%Choices('D', 'P', 'C', 'O', 'F', 'G', 'S') Bracken: specifies the taxonomic rank to analyze. Each classification at this specified rank will receive an estimated number of reads belonging to that rank after abundance estimation.[default:
'S']- include_unclassified:
Bool Bracken does not include the unclassified read counts in the feature table. Set this to True to include those regardless.[default:
True]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('bracken')] Reports modified by Bracken.[required]
- taxonomy:
FeatureData[Taxonomy] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
annotate build-kraken-db¶
This method builds Kraken 2 and Bracken databases either (1) from provided DNA sequences to build a custom database, or (2) simply fetches pre-built versions from an online resource.
Citations¶
Wood et al., 2019; Lu et al., 2017
Inputs¶
- seqs:
List[FeatureData[Sequence]] Sequences to be added to the Kraken 2 database.[optional]
Parameters¶
- collection:
Str%Choices('viral', 'minusb', 'standard', 'standard8', 'standard16', 'pluspf', 'pluspf8', 'pluspf16', 'pluspfp', 'pluspfp8', 'pluspfp16', 'eupathdb', 'nt', 'corent', 'gtdb', 'greengenes', 'rdp', 'silva132', 'silva138') Name of the database collection to be fetched. Please check https://
benlangmead .github .io /aws -indexes /k2 for the description of the available options.[optional] - threads:
Int%Range(1, None) Number of threads. Only applicable when building a custom database.[default:
1]- kmer_len:
Int%Range(1, None) K-mer length in bp/aa.[default:
35]- minimizer_len:
Int%Range(1, None) Minimizer length in bp/aa.[default:
31]- minimizer_spaces:
Int%Range(1, None) Number of characters in minimizer that are ignored in comparisons.[default:
7]- no_masking:
Bool Avoid masking low-complexity sequences prior to building; masking requires dustmasker or segmasker to be installed in PATH.[default:
False]- max_db_size:
Int%Range(0, None) Maximum number of bytes for Kraken 2 hash table; if the estimator determines more would normally be needed, the reference library will be downsampled to fit.[default:
0]- use_ftp:
Bool Use FTP for downloading instead of RSYNC.[default:
False]- load_factor:
Float%Range(0, 1) Proportion of the hash table to be populated.[default:
0.7]- fast_build:
Bool Do not require database to be deterministically built when using multiple threads. This is faster, but does introduce variability in minimizer/LCA pairs.[default:
False]- read_len:
List[Int%Range(1, None)] Ideal read lengths to be used while building the Bracken database.[optional]
Outputs¶
annotate inspect-kraken2-db¶
This method generates a report of identical format to those generated by classify_kraken2, with a slightly different interpretation. Instead of reporting the number of inputs classified to a taxon/clade, the report displays the number of minimizers mapped to each taxon/clade.
Citations¶
Inputs¶
- db:
Kraken2DB The Kraken 2 database for which to generate the report.[required]
Parameters¶
Outputs¶
- report:
Kraken2DBReport The report of the supplied database.[required]
annotate kraken2-to-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
SampleData[Kraken2Report] Per-sample Kraken 2 reports.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- table:
FeatureTable[PresenceAbsence] A presence/absence table of selected features. The features are not of even ranks, but will be the most specific rank available.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate kraken2-to-mag-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
FeatureData[Kraken2Report % Properties('mags')] Per-sample Kraken 2 reports.[required]
- outputs:
FeatureData[Kraken2Output % Properties('mags')] Per-sample Kraken 2 output files.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate build-custom-diamond-db¶
Creates an artifact containing a binary DIAMOND database file (ref_db.dmnd) from a protein reference database file in FASTA format.
Citations¶
Inputs¶
- seqs:
FeatureData[ProteinSequence] Protein reference database.[required]
- taxonomy:
ReferenceDB[NCBITaxonomy] Reference taxonomy, needed to provide taxonomy features.[optional]
Parameters¶
- threads:
Int%Range(1, None) Number of CPU threads.[default:
1]- file_buffer_size:
Int%Range(1, None) File buffer size in bytes.[default:
67108864]- ignore_warnings:
Bool Ignore warnings.[default:
False]- no_parse_seqids:
Bool Print raw seqids without parsing.[default:
False]
Outputs¶
- db:
ReferenceDB[Diamond] DIAMOND database.[required]
annotate fetch-eggnog-db¶
Downloads EggNOG reference database using the download_eggnog_data.py script from eggNOG. Here, this script downloads 3 files and stores them in the output artifact. At least 80 GB of storage space is required to run this action.
Citations¶
Outputs¶
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
annotate fetch-diamond-db¶
Downloads Diamond reference database. This action downloads 1 file (ref_db.dmnd). At least 18 GB of storage space is required to run this action.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database.[required]
annotate fetch-eggnog-proteins¶
Downloads eggNOG proteome database. This script downloads 2 files (e5.proteomes.faa and e5.taxid_info.tsv) and creates and artifact with them. At least 18 GB of storage space is required to run this action.
Citations¶
Outputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information.[required]
annotate fetch-ncbi-taxonomy¶
Downloads NCBI reference taxonomy from the NCBI FTP server. The resulting artifact is required by the build-custom-diamond-db action if one wishes to create a Diamond data base with taxonomy features. At least 30 GB of storage space is required to run this action.
Citations¶
National Center for Biotechnology Information (NCBI), n.d.
Outputs¶
- taxonomy:
ReferenceDB[NCBITaxonomy] NCBI reference taxonomy.[required]
annotate build-eggnog-diamond-db¶
Creates a DIAMOND database which contains the protein sequences that belong to the specified taxon.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information (generated through the
fetch-eggnog-proteinsaction).[required]
Parameters¶
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database for the specified taxon.[required]
annotate -eggnog-diamond-search¶
This method performs the steps by which we find our possible target sequences to annotate using the Diamond search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for ortholog hits.[required]
- db:
ReferenceDB[Diamond] Diamond database.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-hmmer-search¶
This method performs the steps by which we find our possible target sequences to annotate using the HMMER search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-feature-table¶
Create an eggnog table.
Inputs¶
- seed_orthologs:
SampleData[Orthologs] Sequence data to be turned into an eggnog feature table.[required]
Outputs¶
- table:
FeatureTable[Frequency] <no description>[required]
annotate -eggnog-annotate¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] <no description>[required]
- db:
ReferenceDB[Eggnog] <no description>[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggNOG database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] <no description>[required]
annotate predict-genes-prodigal¶
Prodigal (PROkaryotic DYnamic programming Gene-finding ALgorithm), a gene prediction algorithm designed for improved gene structure prediction, translation initiation site recognition, and reduced false positives in bacterial and archaeal genomes.
Citations¶
Inputs¶
- seqs:
FeatureData[MAG]|SampleData[MAGs]|SampleData[Contigs] MAGs or contigs for which one wishes to predict genes.[required]
Parameters¶
- translation_table_number:
Str%Choices('1', '2', '3', '4', '5', '6', '9', '10', '11', '12', '13', '14', '15', '16', '21', '22', '23', '24', '25') Translation table to be used to translate genes into sequences of amino acids. See https://
www .ncbi .nlm .nih .gov /Taxonomy /Utils /wprintgc .cgi for reference.[default: '11']- mode:
Str%Choices('single', 'meta') Gene prediction mode. 'single' is suitable for single genome analysis (e.g., MAGs), 'meta' is suitable for metagenome analysis (e.g., contigs from mixed communities).[default:
'single']- closed:
Bool Treat sequences as complete genomes with closed ends. Use this for finished genomes where the sequences represent complete chromosomes or plasmids.[default:
False]- no_shine_dalgarno:
Bool Bypass Shine-Dalgarno trainer and use a more generic model. Useful for virus, phage, or plasmid sequences that may not follow standard prokaryotic gene patterns.[default:
False]- mask:
Bool Treat runs of N as masked sequence. Useful for assemblies that contain gap regions represented by stretches of N nucleotides.[default:
False]
Outputs¶
- loci:
GenomeData[Loci] Gene coordinates files (one per MAG or sample) listing the location of each predicted gene as well as some additional scoring information.[required]
- genes:
GenomeData[Genes] Fasta files (one per MAG or sample) with the nucleotide sequences of the predicted genes.[required]
- proteins:
GenomeData[Proteins] Fasta files (one per MAG or sample) with the protein translation of the predicted genes.[required]
annotate fetch-kaiju-db¶
This method fetches the latest Kaiju database from Kaiju's web server.
Citations¶
Parameters¶
- database_type:
Str%Choices('nr', 'nr_euk', 'refseq', 'refseq_ref', 'refseq_nr', 'fungi', 'viruses', 'plasmids', 'progenomes', 'rvdb') Type of database to be downloaded. For more information on available types please see the list on Kaiju's web server: https://
bioinformatics -centre .github .io /kaiju /downloads .html[required]
Outputs¶
- db:
KaijuDB Kaiju database.[required]
annotate -classify-kaiju¶
This method uses Kaiju to perform taxonomic classification of DNA sequence reads or contigs.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate fetch-eggnog-hmmer-db¶
Downloads Profile HMM database for the specified taxon.
Citations¶
Huerta-Cepas et al., 2019; HMMER, 2024
Parameters¶
Outputs¶
- idmap:
EggnogHmmerIdmap%Properties('eggnog') List of protein families in
hmm_db.[required]- hmm_db:
ProfileHMM[MultipleProtein]%Properties('eggnog') Collection of Profile HMMs.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein]%Properties('eggnog') Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins]%Properties('eggnog') Seed alignments for the protein families in
hmm_db.[required]
annotate extract-annotations¶
This method extract a specific annotation from the table generated by EggNOG and calculates its frequencies across all contigs (annotation_counts_per_contig) and MAGs (annotation_counts_per_genome).
Inputs¶
- ortholog_annotations:
GenomeData[NOG] Ortholog annotations.[required]
Parameters¶
- annotation:
Str%Choices('cog', 'caz', 'kegg_ko', 'kegg_pathway', 'kegg_reaction', 'kegg_module', 'brite', 'ec') Annotation to extract.[required]
- max_evalue:
Float%Range(0, None) <no description>[default:
1.0]- min_score:
Float%Range(0, None) <no description>[default:
0.0]
Outputs¶
- annotation_counts_per_genome:
FeatureTable[Frequency] Feature table with frequency of each annotation per genome.[required]
- annotation_map:
FeatureMap[FunctionToContigs] Feature map with function to contigs mapping.[required]
- annotation_counts_per_contig:
FeatureTable[Frequency] Feature table with frequency of each annotation per contig.[required]
annotate -filter-kraken2-reports-by-abundance¶
Filters kraken2 reports on a per-taxon basis by relative abundance (relative frequency). Useful for removing suspected spurious classifications.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter by relative abundance.[required]
Parameters¶
- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report.[required]
- remove_empty:
Bool If True, reports with only unclassified reads remaining will be removed from the filtered data.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The relative abundance-filtered kraken2 reports[required]
annotate -filter-kraken2-results-by-metadata¶
Filter Kraken2 reports and outputs based on metadata or remove reports with 100% unclassified reads.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The Kraken reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The Kraken outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] <no description>[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] <no description>[required]
annotate -merge-kraken2-results¶
Merge multiple kraken2 reports and outputs such that the results contain a union of the samples represented in the inputs. If sample IDs overlap across the inputs, these reports and outputs will be processed into a single report or output per sample ID.
Inputs¶
- reports:
List[SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')]] The kraken2 reports to merge. Only reports with the same sample ID are merged into one report.[required]
- outputs:
List[SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')]] The kraken2 outputs to merge. Only outputs with the same sample ID are merged into one output.[required]
Outputs¶
- merged_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The merged kraken2 reports.[required]
- merged_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The merged kraken2 outputs.[required]
annotate -align-outputs-with-reports¶
Inputs¶
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to align with the filtered reports.[required]
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
Outputs¶
- aligned_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The report-aligned filtered kraken2 outputs.[required]
annotate -filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
annotate -visualize-collapsed-contigs¶
Generates a visualization showing histograms of contig abundance distributions per taxon per sample from the original (pre-collapsed) table.
Inputs¶
- table:
FeatureTable[Frequency] Original feature table with contig IDs as feature IDs (before collapsing).[required]
- collapsed_table:
FeatureTable[Frequency] Feature table with contig IDs collapsed to taxonomy IDs.[required]
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings instead of IDs.[optional]
Outputs¶
- visualization:
Visualization <no description>[required]
annotate classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
List[SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]]|List[SampleData[Contigs]]|List[FeatureData[MAG]]|List[SampleData[MAGs]] Sequences to be classified. Single-/paired-end reads,contigs, or assembled MAGs, can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate map-taxonomy-to-contigs¶
Maps contig IDs to their full taxonomy strings based on Kraken 2 classifications. This action processes Kraken 2 reports and outputs produced from contig sequences to create a taxonomy mapping where each contig ID is associated with its assigned taxonomy.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('contigs')] Per-sample Kraken 2 reports for contigs.[required]
- outputs:
SampleData[Kraken2Output % Properties('contigs')] Per-sample Kraken 2 output files for contigs.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- feature_map:
FeatureMap[TaxonomyToContigs] Taxonomy assignments for contigs. Each contig ID is mapped to its full taxonomy string based on Kraken2 classifications. Unclassified contigs are assigned 'd__Unclassified'. Assignments below the provided coverage threshold will be treated as 'd__Unclassified'.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate collapse-contigs¶
This action collapses contig abundances based on their taxonomy assignments. Contig abundances for contigs with the same taxonomy assignment will be averaged.
Inputs¶
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- table:
FeatureTable[Frequency] Table of contig abundances per sample. Feature IDs will be replaced with taxonomy strings.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings in the visualization instead of taxonomy IDs.[optional]
Outputs¶
- collapsed_table:
FeatureTable[Frequency] Contig abundance table with contig IDs replaced by their taxonomy strings. Contigs with the same taxonomy assignment will have their abundances averaged.[required]
- visualization:
Visualization Interactive visualization showing histograms of contig abundance distributions per taxon per sample.[required]
annotate search-orthologs-diamond¶
Use Diamond and eggNOG to align contig or MAG sequences against the Diamond database.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits using the Diamond Database[required]
- db:
ReferenceDB[Diamond] The filepath to an artifact containing the Diamond database[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate search-orthologs-hmmer¶
This method uses HMMER to find possible target sequences to annotate with eggNOG-mapper.
Citations¶
HMMER, 2024; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of profile HMMs in binary format and indexed.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate map-eggnog¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] BLAST6-like table(s) describing the identified orthologs.[required]
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggnog database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] Annotated hits.[required]
annotate classify-kaiju¶
This method uses Kaiju to perform taxonomic classification.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate filter-kraken2-results¶
Filter kraken2 reports and outputs by sample metadata, and/or filter classified taxa by relative abundance.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report. If a taxon is filtered from the report, its associated read counts are removed entirely from the report (i.e., the subtraction of those counts is propagated to parent taxonomic groupings).[optional]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The filtered kraken2 outputs.[required]
annotate filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[default:
1]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
MOdular SHotgun metagenome Pipelines with Integrated provenance Tracking: QIIME 2 plugin gor metagenome analysis withtools for genome binning and functional annotation.
- version:
2026.7.0 - website: https://
github .com /bokulich -lab /q2 -annotate - user support:
- Please post to the QIIME 2 forum for help with this plugin: https://
forum .qiime2 .org
Actions¶
| Name | Type | Short Description |
|---|---|---|
| -classify-kraken2 | method | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| collate-kraken2-reports | method | Collate kraken2 reports. |
| collate-kraken2-outputs | method | Collate kraken2 outputs. |
| estimate-bracken | method | Perform read abundance re-estimation using Bracken. |
| build-kraken-db | method | Build Kraken 2 database. |
| inspect-kraken2-db | method | Inspect a Kraken 2 database. |
| kraken2-to-features | method | Select downstream features from Kraken 2. |
| kraken2-to-mag-features | method | Select downstream MAG features from Kraken 2. |
| build-custom-diamond-db | method | Create a DIAMOND formatted reference database from a FASTA input file. |
| fetch-eggnog-db | method | Fetch the databases necessary to run the eggnog-annotate action. |
| fetch-diamond-db | method | Fetch the complete Diamond database necessary to run the eggnog-diamond-search action. |
| fetch-eggnog-proteins | method | Fetch the databases necessary to run the build-eggnog-diamond-db action. |
| fetch-ncbi-taxonomy | method | Fetch NCBI reference taxonomy. |
| build-eggnog-diamond-db | method | Create a DIAMOND formatted reference database for the specified taxon. |
| -eggnog-diamond-search | method | Run eggNOG search using Diamond aligner. |
| -eggnog-hmmer-search | method | Run eggNOG search using HMMER aligner. |
| -eggnog-feature-table | method | Create an eggnog table. |
| -eggnog-annotate | method | Annotate orthologs against eggNOG database. |
| predict-genes-prodigal | method | Predict gene sequences from MAGs or contigs using Prodigal. |
| fetch-kaiju-db | method | Fetch Kaiju database. |
| -classify-kaiju | method | Classify sequences using Kaiju. |
| fetch-eggnog-hmmer-db | method | Fetch the taxon specific database necessary to run the eggnog-hmmer-search action. |
| extract-annotations | method | Extract annotation frequencies from all annotations. |
| -filter-kraken2-reports-by-abundance | method | Filter kraken2 reports by relative abundance. |
| -filter-kraken2-results-by-metadata | method | Filter Kraken2 reports and outputs. |
| -merge-kraken2-results | method | Merge kraken2 reports and outputs. |
| -align-outputs-with-reports | method | Align unfiltered kraken2 outputs with filtered kraken2 reports. |
| -filter-reads-kraken2 | method | Filter Kraken2-classified reads by taxonomy. |
| -visualize-collapsed-contigs | visualizer | Visualize collapsed contig abundances |
| classify-kraken2 | pipeline | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| map-taxonomy-to-contigs | pipeline | Map contig IDs to taxonomy strings from Kraken 2. |
| collapse-contigs | pipeline | Collapse the contig abundances based on taxonomy. |
| search-orthologs-diamond | pipeline | Run eggNOG search using diamond aligner. |
| search-orthologs-hmmer | pipeline | Run eggNOG search using HMMER aligner. |
| map-eggnog | pipeline | Annotate orthologs against eggNOG database. |
| classify-kaiju | pipeline | Classify sequences using Kaiju. |
| filter-kraken2-results | pipeline | Filter kraken2 reports and outputs by metadata and abundance. |
| filter-reads-kraken2 | pipeline | Filter Kraken2-classified reads by taxonomy. |
Artifact Classes¶
EggnogHmmerIdmap |
Formats¶
EggnogHmmerIdmapFileFmt |
EggnogHmmerIdmapDirectoryFmt |
annotate -classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs]|FeatureData[MAG]|SampleData[MAGs] Sequences to be classified. Single-/paired-end reads, contigs, or assembled MAGs can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate collate-kraken2-reports¶
Collates kraken2 reports.
Inputs¶
- reports:
List[SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_reports:
SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate collate-kraken2-outputs¶
Collates kraken2 outputs.
Inputs¶
- outputs:
List[SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_outputs:
SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate estimate-bracken¶
This method uses Bracken to re-estimate read abundances. Only available on Linux platforms.
Citations¶
Inputs¶
- kraken2_reports:
SampleData[Kraken2Report % Properties('reads')] Reports produced by Kraken2.[required]
- db:
BrackenDB Bracken database.[required]
Parameters¶
- threshold:
Int%Range(0, None) Bracken: number of reads required PRIOR to abundance estimation to perform re-estimation.[default:
0]- read_len:
Int%Range(0, None) Bracken: the ideal length of reads in your sample. For paired end data (e.g., 2x150) this should be set to the length of the single-end reads (e.g., 150).[default:
100]- level:
Str%Choices('D', 'P', 'C', 'O', 'F', 'G', 'S') Bracken: specifies the taxonomic rank to analyze. Each classification at this specified rank will receive an estimated number of reads belonging to that rank after abundance estimation.[default:
'S']- include_unclassified:
Bool Bracken does not include the unclassified read counts in the feature table. Set this to True to include those regardless.[default:
True]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('bracken')] Reports modified by Bracken.[required]
- taxonomy:
FeatureData[Taxonomy] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
annotate build-kraken-db¶
This method builds Kraken 2 and Bracken databases either (1) from provided DNA sequences to build a custom database, or (2) simply fetches pre-built versions from an online resource.
Citations¶
Wood et al., 2019; Lu et al., 2017
Inputs¶
- seqs:
List[FeatureData[Sequence]] Sequences to be added to the Kraken 2 database.[optional]
Parameters¶
- collection:
Str%Choices('viral', 'minusb', 'standard', 'standard8', 'standard16', 'pluspf', 'pluspf8', 'pluspf16', 'pluspfp', 'pluspfp8', 'pluspfp16', 'eupathdb', 'nt', 'corent', 'gtdb', 'greengenes', 'rdp', 'silva132', 'silva138') Name of the database collection to be fetched. Please check https://
benlangmead .github .io /aws -indexes /k2 for the description of the available options.[optional] - threads:
Int%Range(1, None) Number of threads. Only applicable when building a custom database.[default:
1]- kmer_len:
Int%Range(1, None) K-mer length in bp/aa.[default:
35]- minimizer_len:
Int%Range(1, None) Minimizer length in bp/aa.[default:
31]- minimizer_spaces:
Int%Range(1, None) Number of characters in minimizer that are ignored in comparisons.[default:
7]- no_masking:
Bool Avoid masking low-complexity sequences prior to building; masking requires dustmasker or segmasker to be installed in PATH.[default:
False]- max_db_size:
Int%Range(0, None) Maximum number of bytes for Kraken 2 hash table; if the estimator determines more would normally be needed, the reference library will be downsampled to fit.[default:
0]- use_ftp:
Bool Use FTP for downloading instead of RSYNC.[default:
False]- load_factor:
Float%Range(0, 1) Proportion of the hash table to be populated.[default:
0.7]- fast_build:
Bool Do not require database to be deterministically built when using multiple threads. This is faster, but does introduce variability in minimizer/LCA pairs.[default:
False]- read_len:
List[Int%Range(1, None)] Ideal read lengths to be used while building the Bracken database.[optional]
Outputs¶
annotate inspect-kraken2-db¶
This method generates a report of identical format to those generated by classify_kraken2, with a slightly different interpretation. Instead of reporting the number of inputs classified to a taxon/clade, the report displays the number of minimizers mapped to each taxon/clade.
Citations¶
Inputs¶
- db:
Kraken2DB The Kraken 2 database for which to generate the report.[required]
Parameters¶
Outputs¶
- report:
Kraken2DBReport The report of the supplied database.[required]
annotate kraken2-to-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
SampleData[Kraken2Report] Per-sample Kraken 2 reports.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- table:
FeatureTable[PresenceAbsence] A presence/absence table of selected features. The features are not of even ranks, but will be the most specific rank available.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate kraken2-to-mag-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
FeatureData[Kraken2Report % Properties('mags')] Per-sample Kraken 2 reports.[required]
- outputs:
FeatureData[Kraken2Output % Properties('mags')] Per-sample Kraken 2 output files.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate build-custom-diamond-db¶
Creates an artifact containing a binary DIAMOND database file (ref_db.dmnd) from a protein reference database file in FASTA format.
Citations¶
Inputs¶
- seqs:
FeatureData[ProteinSequence] Protein reference database.[required]
- taxonomy:
ReferenceDB[NCBITaxonomy] Reference taxonomy, needed to provide taxonomy features.[optional]
Parameters¶
- threads:
Int%Range(1, None) Number of CPU threads.[default:
1]- file_buffer_size:
Int%Range(1, None) File buffer size in bytes.[default:
67108864]- ignore_warnings:
Bool Ignore warnings.[default:
False]- no_parse_seqids:
Bool Print raw seqids without parsing.[default:
False]
Outputs¶
- db:
ReferenceDB[Diamond] DIAMOND database.[required]
annotate fetch-eggnog-db¶
Downloads EggNOG reference database using the download_eggnog_data.py script from eggNOG. Here, this script downloads 3 files and stores them in the output artifact. At least 80 GB of storage space is required to run this action.
Citations¶
Outputs¶
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
annotate fetch-diamond-db¶
Downloads Diamond reference database. This action downloads 1 file (ref_db.dmnd). At least 18 GB of storage space is required to run this action.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database.[required]
annotate fetch-eggnog-proteins¶
Downloads eggNOG proteome database. This script downloads 2 files (e5.proteomes.faa and e5.taxid_info.tsv) and creates and artifact with them. At least 18 GB of storage space is required to run this action.
Citations¶
Outputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information.[required]
annotate fetch-ncbi-taxonomy¶
Downloads NCBI reference taxonomy from the NCBI FTP server. The resulting artifact is required by the build-custom-diamond-db action if one wishes to create a Diamond data base with taxonomy features. At least 30 GB of storage space is required to run this action.
Citations¶
National Center for Biotechnology Information (NCBI), n.d.
Outputs¶
- taxonomy:
ReferenceDB[NCBITaxonomy] NCBI reference taxonomy.[required]
annotate build-eggnog-diamond-db¶
Creates a DIAMOND database which contains the protein sequences that belong to the specified taxon.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information (generated through the
fetch-eggnog-proteinsaction).[required]
Parameters¶
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database for the specified taxon.[required]
annotate -eggnog-diamond-search¶
This method performs the steps by which we find our possible target sequences to annotate using the Diamond search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for ortholog hits.[required]
- db:
ReferenceDB[Diamond] Diamond database.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-hmmer-search¶
This method performs the steps by which we find our possible target sequences to annotate using the HMMER search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-feature-table¶
Create an eggnog table.
Inputs¶
- seed_orthologs:
SampleData[Orthologs] Sequence data to be turned into an eggnog feature table.[required]
Outputs¶
- table:
FeatureTable[Frequency] <no description>[required]
annotate -eggnog-annotate¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] <no description>[required]
- db:
ReferenceDB[Eggnog] <no description>[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggNOG database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] <no description>[required]
annotate predict-genes-prodigal¶
Prodigal (PROkaryotic DYnamic programming Gene-finding ALgorithm), a gene prediction algorithm designed for improved gene structure prediction, translation initiation site recognition, and reduced false positives in bacterial and archaeal genomes.
Citations¶
Inputs¶
- seqs:
FeatureData[MAG]|SampleData[MAGs]|SampleData[Contigs] MAGs or contigs for which one wishes to predict genes.[required]
Parameters¶
- translation_table_number:
Str%Choices('1', '2', '3', '4', '5', '6', '9', '10', '11', '12', '13', '14', '15', '16', '21', '22', '23', '24', '25') Translation table to be used to translate genes into sequences of amino acids. See https://
www .ncbi .nlm .nih .gov /Taxonomy /Utils /wprintgc .cgi for reference.[default: '11']- mode:
Str%Choices('single', 'meta') Gene prediction mode. 'single' is suitable for single genome analysis (e.g., MAGs), 'meta' is suitable for metagenome analysis (e.g., contigs from mixed communities).[default:
'single']- closed:
Bool Treat sequences as complete genomes with closed ends. Use this for finished genomes where the sequences represent complete chromosomes or plasmids.[default:
False]- no_shine_dalgarno:
Bool Bypass Shine-Dalgarno trainer and use a more generic model. Useful for virus, phage, or plasmid sequences that may not follow standard prokaryotic gene patterns.[default:
False]- mask:
Bool Treat runs of N as masked sequence. Useful for assemblies that contain gap regions represented by stretches of N nucleotides.[default:
False]
Outputs¶
- loci:
GenomeData[Loci] Gene coordinates files (one per MAG or sample) listing the location of each predicted gene as well as some additional scoring information.[required]
- genes:
GenomeData[Genes] Fasta files (one per MAG or sample) with the nucleotide sequences of the predicted genes.[required]
- proteins:
GenomeData[Proteins] Fasta files (one per MAG or sample) with the protein translation of the predicted genes.[required]
annotate fetch-kaiju-db¶
This method fetches the latest Kaiju database from Kaiju's web server.
Citations¶
Parameters¶
- database_type:
Str%Choices('nr', 'nr_euk', 'refseq', 'refseq_ref', 'refseq_nr', 'fungi', 'viruses', 'plasmids', 'progenomes', 'rvdb') Type of database to be downloaded. For more information on available types please see the list on Kaiju's web server: https://
bioinformatics -centre .github .io /kaiju /downloads .html[required]
Outputs¶
- db:
KaijuDB Kaiju database.[required]
annotate -classify-kaiju¶
This method uses Kaiju to perform taxonomic classification of DNA sequence reads or contigs.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate fetch-eggnog-hmmer-db¶
Downloads Profile HMM database for the specified taxon.
Citations¶
Huerta-Cepas et al., 2019; HMMER, 2024
Parameters¶
Outputs¶
- idmap:
EggnogHmmerIdmap%Properties('eggnog') List of protein families in
hmm_db.[required]- hmm_db:
ProfileHMM[MultipleProtein]%Properties('eggnog') Collection of Profile HMMs.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein]%Properties('eggnog') Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins]%Properties('eggnog') Seed alignments for the protein families in
hmm_db.[required]
annotate extract-annotations¶
This method extract a specific annotation from the table generated by EggNOG and calculates its frequencies across all contigs (annotation_counts_per_contig) and MAGs (annotation_counts_per_genome).
Inputs¶
- ortholog_annotations:
GenomeData[NOG] Ortholog annotations.[required]
Parameters¶
- annotation:
Str%Choices('cog', 'caz', 'kegg_ko', 'kegg_pathway', 'kegg_reaction', 'kegg_module', 'brite', 'ec') Annotation to extract.[required]
- max_evalue:
Float%Range(0, None) <no description>[default:
1.0]- min_score:
Float%Range(0, None) <no description>[default:
0.0]
Outputs¶
- annotation_counts_per_genome:
FeatureTable[Frequency] Feature table with frequency of each annotation per genome.[required]
- annotation_map:
FeatureMap[FunctionToContigs] Feature map with function to contigs mapping.[required]
- annotation_counts_per_contig:
FeatureTable[Frequency] Feature table with frequency of each annotation per contig.[required]
annotate -filter-kraken2-reports-by-abundance¶
Filters kraken2 reports on a per-taxon basis by relative abundance (relative frequency). Useful for removing suspected spurious classifications.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter by relative abundance.[required]
Parameters¶
- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report.[required]
- remove_empty:
Bool If True, reports with only unclassified reads remaining will be removed from the filtered data.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The relative abundance-filtered kraken2 reports[required]
annotate -filter-kraken2-results-by-metadata¶
Filter Kraken2 reports and outputs based on metadata or remove reports with 100% unclassified reads.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The Kraken reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The Kraken outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] <no description>[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] <no description>[required]
annotate -merge-kraken2-results¶
Merge multiple kraken2 reports and outputs such that the results contain a union of the samples represented in the inputs. If sample IDs overlap across the inputs, these reports and outputs will be processed into a single report or output per sample ID.
Inputs¶
- reports:
List[SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')]] The kraken2 reports to merge. Only reports with the same sample ID are merged into one report.[required]
- outputs:
List[SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')]] The kraken2 outputs to merge. Only outputs with the same sample ID are merged into one output.[required]
Outputs¶
- merged_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The merged kraken2 reports.[required]
- merged_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The merged kraken2 outputs.[required]
annotate -align-outputs-with-reports¶
Inputs¶
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to align with the filtered reports.[required]
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
Outputs¶
- aligned_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The report-aligned filtered kraken2 outputs.[required]
annotate -filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
annotate -visualize-collapsed-contigs¶
Generates a visualization showing histograms of contig abundance distributions per taxon per sample from the original (pre-collapsed) table.
Inputs¶
- table:
FeatureTable[Frequency] Original feature table with contig IDs as feature IDs (before collapsing).[required]
- collapsed_table:
FeatureTable[Frequency] Feature table with contig IDs collapsed to taxonomy IDs.[required]
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings instead of IDs.[optional]
Outputs¶
- visualization:
Visualization <no description>[required]
annotate classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
List[SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]]|List[SampleData[Contigs]]|List[FeatureData[MAG]]|List[SampleData[MAGs]] Sequences to be classified. Single-/paired-end reads,contigs, or assembled MAGs, can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate map-taxonomy-to-contigs¶
Maps contig IDs to their full taxonomy strings based on Kraken 2 classifications. This action processes Kraken 2 reports and outputs produced from contig sequences to create a taxonomy mapping where each contig ID is associated with its assigned taxonomy.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('contigs')] Per-sample Kraken 2 reports for contigs.[required]
- outputs:
SampleData[Kraken2Output % Properties('contigs')] Per-sample Kraken 2 output files for contigs.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- feature_map:
FeatureMap[TaxonomyToContigs] Taxonomy assignments for contigs. Each contig ID is mapped to its full taxonomy string based on Kraken2 classifications. Unclassified contigs are assigned 'd__Unclassified'. Assignments below the provided coverage threshold will be treated as 'd__Unclassified'.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate collapse-contigs¶
This action collapses contig abundances based on their taxonomy assignments. Contig abundances for contigs with the same taxonomy assignment will be averaged.
Inputs¶
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- table:
FeatureTable[Frequency] Table of contig abundances per sample. Feature IDs will be replaced with taxonomy strings.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings in the visualization instead of taxonomy IDs.[optional]
Outputs¶
- collapsed_table:
FeatureTable[Frequency] Contig abundance table with contig IDs replaced by their taxonomy strings. Contigs with the same taxonomy assignment will have their abundances averaged.[required]
- visualization:
Visualization Interactive visualization showing histograms of contig abundance distributions per taxon per sample.[required]
annotate search-orthologs-diamond¶
Use Diamond and eggNOG to align contig or MAG sequences against the Diamond database.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits using the Diamond Database[required]
- db:
ReferenceDB[Diamond] The filepath to an artifact containing the Diamond database[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate search-orthologs-hmmer¶
This method uses HMMER to find possible target sequences to annotate with eggNOG-mapper.
Citations¶
HMMER, 2024; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of profile HMMs in binary format and indexed.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate map-eggnog¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] BLAST6-like table(s) describing the identified orthologs.[required]
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggnog database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] Annotated hits.[required]
annotate classify-kaiju¶
This method uses Kaiju to perform taxonomic classification.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate filter-kraken2-results¶
Filter kraken2 reports and outputs by sample metadata, and/or filter classified taxa by relative abundance.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report. If a taxon is filtered from the report, its associated read counts are removed entirely from the report (i.e., the subtraction of those counts is propagated to parent taxonomic groupings).[optional]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The filtered kraken2 outputs.[required]
annotate filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[default:
1]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
MOdular SHotgun metagenome Pipelines with Integrated provenance Tracking: QIIME 2 plugin gor metagenome analysis withtools for genome binning and functional annotation.
- version:
2026.7.0 - website: https://
github .com /bokulich -lab /q2 -annotate - user support:
- Please post to the QIIME 2 forum for help with this plugin: https://
forum .qiime2 .org
Actions¶
| Name | Type | Short Description |
|---|---|---|
| -classify-kraken2 | method | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| collate-kraken2-reports | method | Collate kraken2 reports. |
| collate-kraken2-outputs | method | Collate kraken2 outputs. |
| estimate-bracken | method | Perform read abundance re-estimation using Bracken. |
| build-kraken-db | method | Build Kraken 2 database. |
| inspect-kraken2-db | method | Inspect a Kraken 2 database. |
| kraken2-to-features | method | Select downstream features from Kraken 2. |
| kraken2-to-mag-features | method | Select downstream MAG features from Kraken 2. |
| build-custom-diamond-db | method | Create a DIAMOND formatted reference database from a FASTA input file. |
| fetch-eggnog-db | method | Fetch the databases necessary to run the eggnog-annotate action. |
| fetch-diamond-db | method | Fetch the complete Diamond database necessary to run the eggnog-diamond-search action. |
| fetch-eggnog-proteins | method | Fetch the databases necessary to run the build-eggnog-diamond-db action. |
| fetch-ncbi-taxonomy | method | Fetch NCBI reference taxonomy. |
| build-eggnog-diamond-db | method | Create a DIAMOND formatted reference database for the specified taxon. |
| -eggnog-diamond-search | method | Run eggNOG search using Diamond aligner. |
| -eggnog-hmmer-search | method | Run eggNOG search using HMMER aligner. |
| -eggnog-feature-table | method | Create an eggnog table. |
| -eggnog-annotate | method | Annotate orthologs against eggNOG database. |
| predict-genes-prodigal | method | Predict gene sequences from MAGs or contigs using Prodigal. |
| fetch-kaiju-db | method | Fetch Kaiju database. |
| -classify-kaiju | method | Classify sequences using Kaiju. |
| fetch-eggnog-hmmer-db | method | Fetch the taxon specific database necessary to run the eggnog-hmmer-search action. |
| extract-annotations | method | Extract annotation frequencies from all annotations. |
| -filter-kraken2-reports-by-abundance | method | Filter kraken2 reports by relative abundance. |
| -filter-kraken2-results-by-metadata | method | Filter Kraken2 reports and outputs. |
| -merge-kraken2-results | method | Merge kraken2 reports and outputs. |
| -align-outputs-with-reports | method | Align unfiltered kraken2 outputs with filtered kraken2 reports. |
| -filter-reads-kraken2 | method | Filter Kraken2-classified reads by taxonomy. |
| -visualize-collapsed-contigs | visualizer | Visualize collapsed contig abundances |
| classify-kraken2 | pipeline | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| map-taxonomy-to-contigs | pipeline | Map contig IDs to taxonomy strings from Kraken 2. |
| collapse-contigs | pipeline | Collapse the contig abundances based on taxonomy. |
| search-orthologs-diamond | pipeline | Run eggNOG search using diamond aligner. |
| search-orthologs-hmmer | pipeline | Run eggNOG search using HMMER aligner. |
| map-eggnog | pipeline | Annotate orthologs against eggNOG database. |
| classify-kaiju | pipeline | Classify sequences using Kaiju. |
| filter-kraken2-results | pipeline | Filter kraken2 reports and outputs by metadata and abundance. |
| filter-reads-kraken2 | pipeline | Filter Kraken2-classified reads by taxonomy. |
Artifact Classes¶
EggnogHmmerIdmap |
Formats¶
EggnogHmmerIdmapFileFmt |
EggnogHmmerIdmapDirectoryFmt |
annotate -classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs]|FeatureData[MAG]|SampleData[MAGs] Sequences to be classified. Single-/paired-end reads, contigs, or assembled MAGs can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate collate-kraken2-reports¶
Collates kraken2 reports.
Inputs¶
- reports:
List[SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_reports:
SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate collate-kraken2-outputs¶
Collates kraken2 outputs.
Inputs¶
- outputs:
List[SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_outputs:
SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate estimate-bracken¶
This method uses Bracken to re-estimate read abundances. Only available on Linux platforms.
Citations¶
Inputs¶
- kraken2_reports:
SampleData[Kraken2Report % Properties('reads')] Reports produced by Kraken2.[required]
- db:
BrackenDB Bracken database.[required]
Parameters¶
- threshold:
Int%Range(0, None) Bracken: number of reads required PRIOR to abundance estimation to perform re-estimation.[default:
0]- read_len:
Int%Range(0, None) Bracken: the ideal length of reads in your sample. For paired end data (e.g., 2x150) this should be set to the length of the single-end reads (e.g., 150).[default:
100]- level:
Str%Choices('D', 'P', 'C', 'O', 'F', 'G', 'S') Bracken: specifies the taxonomic rank to analyze. Each classification at this specified rank will receive an estimated number of reads belonging to that rank after abundance estimation.[default:
'S']- include_unclassified:
Bool Bracken does not include the unclassified read counts in the feature table. Set this to True to include those regardless.[default:
True]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('bracken')] Reports modified by Bracken.[required]
- taxonomy:
FeatureData[Taxonomy] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
annotate build-kraken-db¶
This method builds Kraken 2 and Bracken databases either (1) from provided DNA sequences to build a custom database, or (2) simply fetches pre-built versions from an online resource.
Citations¶
Wood et al., 2019; Lu et al., 2017
Inputs¶
- seqs:
List[FeatureData[Sequence]] Sequences to be added to the Kraken 2 database.[optional]
Parameters¶
- collection:
Str%Choices('viral', 'minusb', 'standard', 'standard8', 'standard16', 'pluspf', 'pluspf8', 'pluspf16', 'pluspfp', 'pluspfp8', 'pluspfp16', 'eupathdb', 'nt', 'corent', 'gtdb', 'greengenes', 'rdp', 'silva132', 'silva138') Name of the database collection to be fetched. Please check https://
benlangmead .github .io /aws -indexes /k2 for the description of the available options.[optional] - threads:
Int%Range(1, None) Number of threads. Only applicable when building a custom database.[default:
1]- kmer_len:
Int%Range(1, None) K-mer length in bp/aa.[default:
35]- minimizer_len:
Int%Range(1, None) Minimizer length in bp/aa.[default:
31]- minimizer_spaces:
Int%Range(1, None) Number of characters in minimizer that are ignored in comparisons.[default:
7]- no_masking:
Bool Avoid masking low-complexity sequences prior to building; masking requires dustmasker or segmasker to be installed in PATH.[default:
False]- max_db_size:
Int%Range(0, None) Maximum number of bytes for Kraken 2 hash table; if the estimator determines more would normally be needed, the reference library will be downsampled to fit.[default:
0]- use_ftp:
Bool Use FTP for downloading instead of RSYNC.[default:
False]- load_factor:
Float%Range(0, 1) Proportion of the hash table to be populated.[default:
0.7]- fast_build:
Bool Do not require database to be deterministically built when using multiple threads. This is faster, but does introduce variability in minimizer/LCA pairs.[default:
False]- read_len:
List[Int%Range(1, None)] Ideal read lengths to be used while building the Bracken database.[optional]
Outputs¶
annotate inspect-kraken2-db¶
This method generates a report of identical format to those generated by classify_kraken2, with a slightly different interpretation. Instead of reporting the number of inputs classified to a taxon/clade, the report displays the number of minimizers mapped to each taxon/clade.
Citations¶
Inputs¶
- db:
Kraken2DB The Kraken 2 database for which to generate the report.[required]
Parameters¶
Outputs¶
- report:
Kraken2DBReport The report of the supplied database.[required]
annotate kraken2-to-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
SampleData[Kraken2Report] Per-sample Kraken 2 reports.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- table:
FeatureTable[PresenceAbsence] A presence/absence table of selected features. The features are not of even ranks, but will be the most specific rank available.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate kraken2-to-mag-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
FeatureData[Kraken2Report % Properties('mags')] Per-sample Kraken 2 reports.[required]
- outputs:
FeatureData[Kraken2Output % Properties('mags')] Per-sample Kraken 2 output files.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate build-custom-diamond-db¶
Creates an artifact containing a binary DIAMOND database file (ref_db.dmnd) from a protein reference database file in FASTA format.
Citations¶
Inputs¶
- seqs:
FeatureData[ProteinSequence] Protein reference database.[required]
- taxonomy:
ReferenceDB[NCBITaxonomy] Reference taxonomy, needed to provide taxonomy features.[optional]
Parameters¶
- threads:
Int%Range(1, None) Number of CPU threads.[default:
1]- file_buffer_size:
Int%Range(1, None) File buffer size in bytes.[default:
67108864]- ignore_warnings:
Bool Ignore warnings.[default:
False]- no_parse_seqids:
Bool Print raw seqids without parsing.[default:
False]
Outputs¶
- db:
ReferenceDB[Diamond] DIAMOND database.[required]
annotate fetch-eggnog-db¶
Downloads EggNOG reference database using the download_eggnog_data.py script from eggNOG. Here, this script downloads 3 files and stores them in the output artifact. At least 80 GB of storage space is required to run this action.
Citations¶
Outputs¶
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
annotate fetch-diamond-db¶
Downloads Diamond reference database. This action downloads 1 file (ref_db.dmnd). At least 18 GB of storage space is required to run this action.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database.[required]
annotate fetch-eggnog-proteins¶
Downloads eggNOG proteome database. This script downloads 2 files (e5.proteomes.faa and e5.taxid_info.tsv) and creates and artifact with them. At least 18 GB of storage space is required to run this action.
Citations¶
Outputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information.[required]
annotate fetch-ncbi-taxonomy¶
Downloads NCBI reference taxonomy from the NCBI FTP server. The resulting artifact is required by the build-custom-diamond-db action if one wishes to create a Diamond data base with taxonomy features. At least 30 GB of storage space is required to run this action.
Citations¶
National Center for Biotechnology Information (NCBI), n.d.
Outputs¶
- taxonomy:
ReferenceDB[NCBITaxonomy] NCBI reference taxonomy.[required]
annotate build-eggnog-diamond-db¶
Creates a DIAMOND database which contains the protein sequences that belong to the specified taxon.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information (generated through the
fetch-eggnog-proteinsaction).[required]
Parameters¶
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database for the specified taxon.[required]
annotate -eggnog-diamond-search¶
This method performs the steps by which we find our possible target sequences to annotate using the Diamond search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for ortholog hits.[required]
- db:
ReferenceDB[Diamond] Diamond database.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-hmmer-search¶
This method performs the steps by which we find our possible target sequences to annotate using the HMMER search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-feature-table¶
Create an eggnog table.
Inputs¶
- seed_orthologs:
SampleData[Orthologs] Sequence data to be turned into an eggnog feature table.[required]
Outputs¶
- table:
FeatureTable[Frequency] <no description>[required]
annotate -eggnog-annotate¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] <no description>[required]
- db:
ReferenceDB[Eggnog] <no description>[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggNOG database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] <no description>[required]
annotate predict-genes-prodigal¶
Prodigal (PROkaryotic DYnamic programming Gene-finding ALgorithm), a gene prediction algorithm designed for improved gene structure prediction, translation initiation site recognition, and reduced false positives in bacterial and archaeal genomes.
Citations¶
Inputs¶
- seqs:
FeatureData[MAG]|SampleData[MAGs]|SampleData[Contigs] MAGs or contigs for which one wishes to predict genes.[required]
Parameters¶
- translation_table_number:
Str%Choices('1', '2', '3', '4', '5', '6', '9', '10', '11', '12', '13', '14', '15', '16', '21', '22', '23', '24', '25') Translation table to be used to translate genes into sequences of amino acids. See https://
www .ncbi .nlm .nih .gov /Taxonomy /Utils /wprintgc .cgi for reference.[default: '11']- mode:
Str%Choices('single', 'meta') Gene prediction mode. 'single' is suitable for single genome analysis (e.g., MAGs), 'meta' is suitable for metagenome analysis (e.g., contigs from mixed communities).[default:
'single']- closed:
Bool Treat sequences as complete genomes with closed ends. Use this for finished genomes where the sequences represent complete chromosomes or plasmids.[default:
False]- no_shine_dalgarno:
Bool Bypass Shine-Dalgarno trainer and use a more generic model. Useful for virus, phage, or plasmid sequences that may not follow standard prokaryotic gene patterns.[default:
False]- mask:
Bool Treat runs of N as masked sequence. Useful for assemblies that contain gap regions represented by stretches of N nucleotides.[default:
False]
Outputs¶
- loci:
GenomeData[Loci] Gene coordinates files (one per MAG or sample) listing the location of each predicted gene as well as some additional scoring information.[required]
- genes:
GenomeData[Genes] Fasta files (one per MAG or sample) with the nucleotide sequences of the predicted genes.[required]
- proteins:
GenomeData[Proteins] Fasta files (one per MAG or sample) with the protein translation of the predicted genes.[required]
annotate fetch-kaiju-db¶
This method fetches the latest Kaiju database from Kaiju's web server.
Citations¶
Parameters¶
- database_type:
Str%Choices('nr', 'nr_euk', 'refseq', 'refseq_ref', 'refseq_nr', 'fungi', 'viruses', 'plasmids', 'progenomes', 'rvdb') Type of database to be downloaded. For more information on available types please see the list on Kaiju's web server: https://
bioinformatics -centre .github .io /kaiju /downloads .html[required]
Outputs¶
- db:
KaijuDB Kaiju database.[required]
annotate -classify-kaiju¶
This method uses Kaiju to perform taxonomic classification of DNA sequence reads or contigs.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate fetch-eggnog-hmmer-db¶
Downloads Profile HMM database for the specified taxon.
Citations¶
Huerta-Cepas et al., 2019; HMMER, 2024
Parameters¶
Outputs¶
- idmap:
EggnogHmmerIdmap%Properties('eggnog') List of protein families in
hmm_db.[required]- hmm_db:
ProfileHMM[MultipleProtein]%Properties('eggnog') Collection of Profile HMMs.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein]%Properties('eggnog') Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins]%Properties('eggnog') Seed alignments for the protein families in
hmm_db.[required]
annotate extract-annotations¶
This method extract a specific annotation from the table generated by EggNOG and calculates its frequencies across all contigs (annotation_counts_per_contig) and MAGs (annotation_counts_per_genome).
Inputs¶
- ortholog_annotations:
GenomeData[NOG] Ortholog annotations.[required]
Parameters¶
- annotation:
Str%Choices('cog', 'caz', 'kegg_ko', 'kegg_pathway', 'kegg_reaction', 'kegg_module', 'brite', 'ec') Annotation to extract.[required]
- max_evalue:
Float%Range(0, None) <no description>[default:
1.0]- min_score:
Float%Range(0, None) <no description>[default:
0.0]
Outputs¶
- annotation_counts_per_genome:
FeatureTable[Frequency] Feature table with frequency of each annotation per genome.[required]
- annotation_map:
FeatureMap[FunctionToContigs] Feature map with function to contigs mapping.[required]
- annotation_counts_per_contig:
FeatureTable[Frequency] Feature table with frequency of each annotation per contig.[required]
annotate -filter-kraken2-reports-by-abundance¶
Filters kraken2 reports on a per-taxon basis by relative abundance (relative frequency). Useful for removing suspected spurious classifications.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter by relative abundance.[required]
Parameters¶
- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report.[required]
- remove_empty:
Bool If True, reports with only unclassified reads remaining will be removed from the filtered data.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The relative abundance-filtered kraken2 reports[required]
annotate -filter-kraken2-results-by-metadata¶
Filter Kraken2 reports and outputs based on metadata or remove reports with 100% unclassified reads.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The Kraken reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The Kraken outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] <no description>[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] <no description>[required]
annotate -merge-kraken2-results¶
Merge multiple kraken2 reports and outputs such that the results contain a union of the samples represented in the inputs. If sample IDs overlap across the inputs, these reports and outputs will be processed into a single report or output per sample ID.
Inputs¶
- reports:
List[SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')]] The kraken2 reports to merge. Only reports with the same sample ID are merged into one report.[required]
- outputs:
List[SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')]] The kraken2 outputs to merge. Only outputs with the same sample ID are merged into one output.[required]
Outputs¶
- merged_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The merged kraken2 reports.[required]
- merged_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The merged kraken2 outputs.[required]
annotate -align-outputs-with-reports¶
Inputs¶
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to align with the filtered reports.[required]
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
Outputs¶
- aligned_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The report-aligned filtered kraken2 outputs.[required]
annotate -filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
annotate -visualize-collapsed-contigs¶
Generates a visualization showing histograms of contig abundance distributions per taxon per sample from the original (pre-collapsed) table.
Inputs¶
- table:
FeatureTable[Frequency] Original feature table with contig IDs as feature IDs (before collapsing).[required]
- collapsed_table:
FeatureTable[Frequency] Feature table with contig IDs collapsed to taxonomy IDs.[required]
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings instead of IDs.[optional]
Outputs¶
- visualization:
Visualization <no description>[required]
annotate classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
List[SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]]|List[SampleData[Contigs]]|List[FeatureData[MAG]]|List[SampleData[MAGs]] Sequences to be classified. Single-/paired-end reads,contigs, or assembled MAGs, can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate map-taxonomy-to-contigs¶
Maps contig IDs to their full taxonomy strings based on Kraken 2 classifications. This action processes Kraken 2 reports and outputs produced from contig sequences to create a taxonomy mapping where each contig ID is associated with its assigned taxonomy.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('contigs')] Per-sample Kraken 2 reports for contigs.[required]
- outputs:
SampleData[Kraken2Output % Properties('contigs')] Per-sample Kraken 2 output files for contigs.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- feature_map:
FeatureMap[TaxonomyToContigs] Taxonomy assignments for contigs. Each contig ID is mapped to its full taxonomy string based on Kraken2 classifications. Unclassified contigs are assigned 'd__Unclassified'. Assignments below the provided coverage threshold will be treated as 'd__Unclassified'.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate collapse-contigs¶
This action collapses contig abundances based on their taxonomy assignments. Contig abundances for contigs with the same taxonomy assignment will be averaged.
Inputs¶
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- table:
FeatureTable[Frequency] Table of contig abundances per sample. Feature IDs will be replaced with taxonomy strings.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings in the visualization instead of taxonomy IDs.[optional]
Outputs¶
- collapsed_table:
FeatureTable[Frequency] Contig abundance table with contig IDs replaced by their taxonomy strings. Contigs with the same taxonomy assignment will have their abundances averaged.[required]
- visualization:
Visualization Interactive visualization showing histograms of contig abundance distributions per taxon per sample.[required]
annotate search-orthologs-diamond¶
Use Diamond and eggNOG to align contig or MAG sequences against the Diamond database.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits using the Diamond Database[required]
- db:
ReferenceDB[Diamond] The filepath to an artifact containing the Diamond database[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate search-orthologs-hmmer¶
This method uses HMMER to find possible target sequences to annotate with eggNOG-mapper.
Citations¶
HMMER, 2024; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of profile HMMs in binary format and indexed.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate map-eggnog¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] BLAST6-like table(s) describing the identified orthologs.[required]
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggnog database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] Annotated hits.[required]
annotate classify-kaiju¶
This method uses Kaiju to perform taxonomic classification.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate filter-kraken2-results¶
Filter kraken2 reports and outputs by sample metadata, and/or filter classified taxa by relative abundance.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report. If a taxon is filtered from the report, its associated read counts are removed entirely from the report (i.e., the subtraction of those counts is propagated to parent taxonomic groupings).[optional]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The filtered kraken2 outputs.[required]
annotate filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[default:
1]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
MOdular SHotgun metagenome Pipelines with Integrated provenance Tracking: QIIME 2 plugin gor metagenome analysis withtools for genome binning and functional annotation.
- version:
2026.7.0 - website: https://
github .com /bokulich -lab /q2 -annotate - user support:
- Please post to the QIIME 2 forum for help with this plugin: https://
forum .qiime2 .org
Actions¶
| Name | Type | Short Description |
|---|---|---|
| -classify-kraken2 | method | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| collate-kraken2-reports | method | Collate kraken2 reports. |
| collate-kraken2-outputs | method | Collate kraken2 outputs. |
| estimate-bracken | method | Perform read abundance re-estimation using Bracken. |
| build-kraken-db | method | Build Kraken 2 database. |
| inspect-kraken2-db | method | Inspect a Kraken 2 database. |
| kraken2-to-features | method | Select downstream features from Kraken 2. |
| kraken2-to-mag-features | method | Select downstream MAG features from Kraken 2. |
| build-custom-diamond-db | method | Create a DIAMOND formatted reference database from a FASTA input file. |
| fetch-eggnog-db | method | Fetch the databases necessary to run the eggnog-annotate action. |
| fetch-diamond-db | method | Fetch the complete Diamond database necessary to run the eggnog-diamond-search action. |
| fetch-eggnog-proteins | method | Fetch the databases necessary to run the build-eggnog-diamond-db action. |
| fetch-ncbi-taxonomy | method | Fetch NCBI reference taxonomy. |
| build-eggnog-diamond-db | method | Create a DIAMOND formatted reference database for the specified taxon. |
| -eggnog-diamond-search | method | Run eggNOG search using Diamond aligner. |
| -eggnog-hmmer-search | method | Run eggNOG search using HMMER aligner. |
| -eggnog-feature-table | method | Create an eggnog table. |
| -eggnog-annotate | method | Annotate orthologs against eggNOG database. |
| predict-genes-prodigal | method | Predict gene sequences from MAGs or contigs using Prodigal. |
| fetch-kaiju-db | method | Fetch Kaiju database. |
| -classify-kaiju | method | Classify sequences using Kaiju. |
| fetch-eggnog-hmmer-db | method | Fetch the taxon specific database necessary to run the eggnog-hmmer-search action. |
| extract-annotations | method | Extract annotation frequencies from all annotations. |
| -filter-kraken2-reports-by-abundance | method | Filter kraken2 reports by relative abundance. |
| -filter-kraken2-results-by-metadata | method | Filter Kraken2 reports and outputs. |
| -merge-kraken2-results | method | Merge kraken2 reports and outputs. |
| -align-outputs-with-reports | method | Align unfiltered kraken2 outputs with filtered kraken2 reports. |
| -filter-reads-kraken2 | method | Filter Kraken2-classified reads by taxonomy. |
| -visualize-collapsed-contigs | visualizer | Visualize collapsed contig abundances |
| classify-kraken2 | pipeline | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| map-taxonomy-to-contigs | pipeline | Map contig IDs to taxonomy strings from Kraken 2. |
| collapse-contigs | pipeline | Collapse the contig abundances based on taxonomy. |
| search-orthologs-diamond | pipeline | Run eggNOG search using diamond aligner. |
| search-orthologs-hmmer | pipeline | Run eggNOG search using HMMER aligner. |
| map-eggnog | pipeline | Annotate orthologs against eggNOG database. |
| classify-kaiju | pipeline | Classify sequences using Kaiju. |
| filter-kraken2-results | pipeline | Filter kraken2 reports and outputs by metadata and abundance. |
| filter-reads-kraken2 | pipeline | Filter Kraken2-classified reads by taxonomy. |
Artifact Classes¶
EggnogHmmerIdmap |
Formats¶
EggnogHmmerIdmapFileFmt |
EggnogHmmerIdmapDirectoryFmt |
annotate -classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs]|FeatureData[MAG]|SampleData[MAGs] Sequences to be classified. Single-/paired-end reads, contigs, or assembled MAGs can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate collate-kraken2-reports¶
Collates kraken2 reports.
Inputs¶
- reports:
List[SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_reports:
SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate collate-kraken2-outputs¶
Collates kraken2 outputs.
Inputs¶
- outputs:
List[SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_outputs:
SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate estimate-bracken¶
This method uses Bracken to re-estimate read abundances. Only available on Linux platforms.
Citations¶
Inputs¶
- kraken2_reports:
SampleData[Kraken2Report % Properties('reads')] Reports produced by Kraken2.[required]
- db:
BrackenDB Bracken database.[required]
Parameters¶
- threshold:
Int%Range(0, None) Bracken: number of reads required PRIOR to abundance estimation to perform re-estimation.[default:
0]- read_len:
Int%Range(0, None) Bracken: the ideal length of reads in your sample. For paired end data (e.g., 2x150) this should be set to the length of the single-end reads (e.g., 150).[default:
100]- level:
Str%Choices('D', 'P', 'C', 'O', 'F', 'G', 'S') Bracken: specifies the taxonomic rank to analyze. Each classification at this specified rank will receive an estimated number of reads belonging to that rank after abundance estimation.[default:
'S']- include_unclassified:
Bool Bracken does not include the unclassified read counts in the feature table. Set this to True to include those regardless.[default:
True]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('bracken')] Reports modified by Bracken.[required]
- taxonomy:
FeatureData[Taxonomy] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
annotate build-kraken-db¶
This method builds Kraken 2 and Bracken databases either (1) from provided DNA sequences to build a custom database, or (2) simply fetches pre-built versions from an online resource.
Citations¶
Wood et al., 2019; Lu et al., 2017
Inputs¶
- seqs:
List[FeatureData[Sequence]] Sequences to be added to the Kraken 2 database.[optional]
Parameters¶
- collection:
Str%Choices('viral', 'minusb', 'standard', 'standard8', 'standard16', 'pluspf', 'pluspf8', 'pluspf16', 'pluspfp', 'pluspfp8', 'pluspfp16', 'eupathdb', 'nt', 'corent', 'gtdb', 'greengenes', 'rdp', 'silva132', 'silva138') Name of the database collection to be fetched. Please check https://
benlangmead .github .io /aws -indexes /k2 for the description of the available options.[optional] - threads:
Int%Range(1, None) Number of threads. Only applicable when building a custom database.[default:
1]- kmer_len:
Int%Range(1, None) K-mer length in bp/aa.[default:
35]- minimizer_len:
Int%Range(1, None) Minimizer length in bp/aa.[default:
31]- minimizer_spaces:
Int%Range(1, None) Number of characters in minimizer that are ignored in comparisons.[default:
7]- no_masking:
Bool Avoid masking low-complexity sequences prior to building; masking requires dustmasker or segmasker to be installed in PATH.[default:
False]- max_db_size:
Int%Range(0, None) Maximum number of bytes for Kraken 2 hash table; if the estimator determines more would normally be needed, the reference library will be downsampled to fit.[default:
0]- use_ftp:
Bool Use FTP for downloading instead of RSYNC.[default:
False]- load_factor:
Float%Range(0, 1) Proportion of the hash table to be populated.[default:
0.7]- fast_build:
Bool Do not require database to be deterministically built when using multiple threads. This is faster, but does introduce variability in minimizer/LCA pairs.[default:
False]- read_len:
List[Int%Range(1, None)] Ideal read lengths to be used while building the Bracken database.[optional]
Outputs¶
annotate inspect-kraken2-db¶
This method generates a report of identical format to those generated by classify_kraken2, with a slightly different interpretation. Instead of reporting the number of inputs classified to a taxon/clade, the report displays the number of minimizers mapped to each taxon/clade.
Citations¶
Inputs¶
- db:
Kraken2DB The Kraken 2 database for which to generate the report.[required]
Parameters¶
Outputs¶
- report:
Kraken2DBReport The report of the supplied database.[required]
annotate kraken2-to-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
SampleData[Kraken2Report] Per-sample Kraken 2 reports.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- table:
FeatureTable[PresenceAbsence] A presence/absence table of selected features. The features are not of even ranks, but will be the most specific rank available.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate kraken2-to-mag-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
FeatureData[Kraken2Report % Properties('mags')] Per-sample Kraken 2 reports.[required]
- outputs:
FeatureData[Kraken2Output % Properties('mags')] Per-sample Kraken 2 output files.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate build-custom-diamond-db¶
Creates an artifact containing a binary DIAMOND database file (ref_db.dmnd) from a protein reference database file in FASTA format.
Citations¶
Inputs¶
- seqs:
FeatureData[ProteinSequence] Protein reference database.[required]
- taxonomy:
ReferenceDB[NCBITaxonomy] Reference taxonomy, needed to provide taxonomy features.[optional]
Parameters¶
- threads:
Int%Range(1, None) Number of CPU threads.[default:
1]- file_buffer_size:
Int%Range(1, None) File buffer size in bytes.[default:
67108864]- ignore_warnings:
Bool Ignore warnings.[default:
False]- no_parse_seqids:
Bool Print raw seqids without parsing.[default:
False]
Outputs¶
- db:
ReferenceDB[Diamond] DIAMOND database.[required]
annotate fetch-eggnog-db¶
Downloads EggNOG reference database using the download_eggnog_data.py script from eggNOG. Here, this script downloads 3 files and stores them in the output artifact. At least 80 GB of storage space is required to run this action.
Citations¶
Outputs¶
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
annotate fetch-diamond-db¶
Downloads Diamond reference database. This action downloads 1 file (ref_db.dmnd). At least 18 GB of storage space is required to run this action.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database.[required]
annotate fetch-eggnog-proteins¶
Downloads eggNOG proteome database. This script downloads 2 files (e5.proteomes.faa and e5.taxid_info.tsv) and creates and artifact with them. At least 18 GB of storage space is required to run this action.
Citations¶
Outputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information.[required]
annotate fetch-ncbi-taxonomy¶
Downloads NCBI reference taxonomy from the NCBI FTP server. The resulting artifact is required by the build-custom-diamond-db action if one wishes to create a Diamond data base with taxonomy features. At least 30 GB of storage space is required to run this action.
Citations¶
National Center for Biotechnology Information (NCBI), n.d.
Outputs¶
- taxonomy:
ReferenceDB[NCBITaxonomy] NCBI reference taxonomy.[required]
annotate build-eggnog-diamond-db¶
Creates a DIAMOND database which contains the protein sequences that belong to the specified taxon.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information (generated through the
fetch-eggnog-proteinsaction).[required]
Parameters¶
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database for the specified taxon.[required]
annotate -eggnog-diamond-search¶
This method performs the steps by which we find our possible target sequences to annotate using the Diamond search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for ortholog hits.[required]
- db:
ReferenceDB[Diamond] Diamond database.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-hmmer-search¶
This method performs the steps by which we find our possible target sequences to annotate using the HMMER search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-feature-table¶
Create an eggnog table.
Inputs¶
- seed_orthologs:
SampleData[Orthologs] Sequence data to be turned into an eggnog feature table.[required]
Outputs¶
- table:
FeatureTable[Frequency] <no description>[required]
annotate -eggnog-annotate¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] <no description>[required]
- db:
ReferenceDB[Eggnog] <no description>[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggNOG database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] <no description>[required]
annotate predict-genes-prodigal¶
Prodigal (PROkaryotic DYnamic programming Gene-finding ALgorithm), a gene prediction algorithm designed for improved gene structure prediction, translation initiation site recognition, and reduced false positives in bacterial and archaeal genomes.
Citations¶
Inputs¶
- seqs:
FeatureData[MAG]|SampleData[MAGs]|SampleData[Contigs] MAGs or contigs for which one wishes to predict genes.[required]
Parameters¶
- translation_table_number:
Str%Choices('1', '2', '3', '4', '5', '6', '9', '10', '11', '12', '13', '14', '15', '16', '21', '22', '23', '24', '25') Translation table to be used to translate genes into sequences of amino acids. See https://
www .ncbi .nlm .nih .gov /Taxonomy /Utils /wprintgc .cgi for reference.[default: '11']- mode:
Str%Choices('single', 'meta') Gene prediction mode. 'single' is suitable for single genome analysis (e.g., MAGs), 'meta' is suitable for metagenome analysis (e.g., contigs from mixed communities).[default:
'single']- closed:
Bool Treat sequences as complete genomes with closed ends. Use this for finished genomes where the sequences represent complete chromosomes or plasmids.[default:
False]- no_shine_dalgarno:
Bool Bypass Shine-Dalgarno trainer and use a more generic model. Useful for virus, phage, or plasmid sequences that may not follow standard prokaryotic gene patterns.[default:
False]- mask:
Bool Treat runs of N as masked sequence. Useful for assemblies that contain gap regions represented by stretches of N nucleotides.[default:
False]
Outputs¶
- loci:
GenomeData[Loci] Gene coordinates files (one per MAG or sample) listing the location of each predicted gene as well as some additional scoring information.[required]
- genes:
GenomeData[Genes] Fasta files (one per MAG or sample) with the nucleotide sequences of the predicted genes.[required]
- proteins:
GenomeData[Proteins] Fasta files (one per MAG or sample) with the protein translation of the predicted genes.[required]
annotate fetch-kaiju-db¶
This method fetches the latest Kaiju database from Kaiju's web server.
Citations¶
Parameters¶
- database_type:
Str%Choices('nr', 'nr_euk', 'refseq', 'refseq_ref', 'refseq_nr', 'fungi', 'viruses', 'plasmids', 'progenomes', 'rvdb') Type of database to be downloaded. For more information on available types please see the list on Kaiju's web server: https://
bioinformatics -centre .github .io /kaiju /downloads .html[required]
Outputs¶
- db:
KaijuDB Kaiju database.[required]
annotate -classify-kaiju¶
This method uses Kaiju to perform taxonomic classification of DNA sequence reads or contigs.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate fetch-eggnog-hmmer-db¶
Downloads Profile HMM database for the specified taxon.
Citations¶
Huerta-Cepas et al., 2019; HMMER, 2024
Parameters¶
Outputs¶
- idmap:
EggnogHmmerIdmap%Properties('eggnog') List of protein families in
hmm_db.[required]- hmm_db:
ProfileHMM[MultipleProtein]%Properties('eggnog') Collection of Profile HMMs.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein]%Properties('eggnog') Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins]%Properties('eggnog') Seed alignments for the protein families in
hmm_db.[required]
annotate extract-annotations¶
This method extract a specific annotation from the table generated by EggNOG and calculates its frequencies across all contigs (annotation_counts_per_contig) and MAGs (annotation_counts_per_genome).
Inputs¶
- ortholog_annotations:
GenomeData[NOG] Ortholog annotations.[required]
Parameters¶
- annotation:
Str%Choices('cog', 'caz', 'kegg_ko', 'kegg_pathway', 'kegg_reaction', 'kegg_module', 'brite', 'ec') Annotation to extract.[required]
- max_evalue:
Float%Range(0, None) <no description>[default:
1.0]- min_score:
Float%Range(0, None) <no description>[default:
0.0]
Outputs¶
- annotation_counts_per_genome:
FeatureTable[Frequency] Feature table with frequency of each annotation per genome.[required]
- annotation_map:
FeatureMap[FunctionToContigs] Feature map with function to contigs mapping.[required]
- annotation_counts_per_contig:
FeatureTable[Frequency] Feature table with frequency of each annotation per contig.[required]
annotate -filter-kraken2-reports-by-abundance¶
Filters kraken2 reports on a per-taxon basis by relative abundance (relative frequency). Useful for removing suspected spurious classifications.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter by relative abundance.[required]
Parameters¶
- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report.[required]
- remove_empty:
Bool If True, reports with only unclassified reads remaining will be removed from the filtered data.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The relative abundance-filtered kraken2 reports[required]
annotate -filter-kraken2-results-by-metadata¶
Filter Kraken2 reports and outputs based on metadata or remove reports with 100% unclassified reads.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The Kraken reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The Kraken outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] <no description>[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] <no description>[required]
annotate -merge-kraken2-results¶
Merge multiple kraken2 reports and outputs such that the results contain a union of the samples represented in the inputs. If sample IDs overlap across the inputs, these reports and outputs will be processed into a single report or output per sample ID.
Inputs¶
- reports:
List[SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')]] The kraken2 reports to merge. Only reports with the same sample ID are merged into one report.[required]
- outputs:
List[SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')]] The kraken2 outputs to merge. Only outputs with the same sample ID are merged into one output.[required]
Outputs¶
- merged_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The merged kraken2 reports.[required]
- merged_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The merged kraken2 outputs.[required]
annotate -align-outputs-with-reports¶
Inputs¶
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to align with the filtered reports.[required]
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
Outputs¶
- aligned_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The report-aligned filtered kraken2 outputs.[required]
annotate -filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
annotate -visualize-collapsed-contigs¶
Generates a visualization showing histograms of contig abundance distributions per taxon per sample from the original (pre-collapsed) table.
Inputs¶
- table:
FeatureTable[Frequency] Original feature table with contig IDs as feature IDs (before collapsing).[required]
- collapsed_table:
FeatureTable[Frequency] Feature table with contig IDs collapsed to taxonomy IDs.[required]
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings instead of IDs.[optional]
Outputs¶
- visualization:
Visualization <no description>[required]
annotate classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
List[SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]]|List[SampleData[Contigs]]|List[FeatureData[MAG]]|List[SampleData[MAGs]] Sequences to be classified. Single-/paired-end reads,contigs, or assembled MAGs, can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate map-taxonomy-to-contigs¶
Maps contig IDs to their full taxonomy strings based on Kraken 2 classifications. This action processes Kraken 2 reports and outputs produced from contig sequences to create a taxonomy mapping where each contig ID is associated with its assigned taxonomy.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('contigs')] Per-sample Kraken 2 reports for contigs.[required]
- outputs:
SampleData[Kraken2Output % Properties('contigs')] Per-sample Kraken 2 output files for contigs.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- feature_map:
FeatureMap[TaxonomyToContigs] Taxonomy assignments for contigs. Each contig ID is mapped to its full taxonomy string based on Kraken2 classifications. Unclassified contigs are assigned 'd__Unclassified'. Assignments below the provided coverage threshold will be treated as 'd__Unclassified'.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate collapse-contigs¶
This action collapses contig abundances based on their taxonomy assignments. Contig abundances for contigs with the same taxonomy assignment will be averaged.
Inputs¶
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- table:
FeatureTable[Frequency] Table of contig abundances per sample. Feature IDs will be replaced with taxonomy strings.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings in the visualization instead of taxonomy IDs.[optional]
Outputs¶
- collapsed_table:
FeatureTable[Frequency] Contig abundance table with contig IDs replaced by their taxonomy strings. Contigs with the same taxonomy assignment will have their abundances averaged.[required]
- visualization:
Visualization Interactive visualization showing histograms of contig abundance distributions per taxon per sample.[required]
annotate search-orthologs-diamond¶
Use Diamond and eggNOG to align contig or MAG sequences against the Diamond database.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits using the Diamond Database[required]
- db:
ReferenceDB[Diamond] The filepath to an artifact containing the Diamond database[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate search-orthologs-hmmer¶
This method uses HMMER to find possible target sequences to annotate with eggNOG-mapper.
Citations¶
HMMER, 2024; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of profile HMMs in binary format and indexed.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate map-eggnog¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] BLAST6-like table(s) describing the identified orthologs.[required]
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggnog database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] Annotated hits.[required]
annotate classify-kaiju¶
This method uses Kaiju to perform taxonomic classification.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate filter-kraken2-results¶
Filter kraken2 reports and outputs by sample metadata, and/or filter classified taxa by relative abundance.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report. If a taxon is filtered from the report, its associated read counts are removed entirely from the report (i.e., the subtraction of those counts is propagated to parent taxonomic groupings).[optional]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The filtered kraken2 outputs.[required]
annotate filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[default:
1]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
MOdular SHotgun metagenome Pipelines with Integrated provenance Tracking: QIIME 2 plugin gor metagenome analysis withtools for genome binning and functional annotation.
- version:
2026.7.0 - website: https://
github .com /bokulich -lab /q2 -annotate - user support:
- Please post to the QIIME 2 forum for help with this plugin: https://
forum .qiime2 .org
Actions¶
| Name | Type | Short Description |
|---|---|---|
| -classify-kraken2 | method | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| collate-kraken2-reports | method | Collate kraken2 reports. |
| collate-kraken2-outputs | method | Collate kraken2 outputs. |
| estimate-bracken | method | Perform read abundance re-estimation using Bracken. |
| build-kraken-db | method | Build Kraken 2 database. |
| inspect-kraken2-db | method | Inspect a Kraken 2 database. |
| kraken2-to-features | method | Select downstream features from Kraken 2. |
| kraken2-to-mag-features | method | Select downstream MAG features from Kraken 2. |
| build-custom-diamond-db | method | Create a DIAMOND formatted reference database from a FASTA input file. |
| fetch-eggnog-db | method | Fetch the databases necessary to run the eggnog-annotate action. |
| fetch-diamond-db | method | Fetch the complete Diamond database necessary to run the eggnog-diamond-search action. |
| fetch-eggnog-proteins | method | Fetch the databases necessary to run the build-eggnog-diamond-db action. |
| fetch-ncbi-taxonomy | method | Fetch NCBI reference taxonomy. |
| build-eggnog-diamond-db | method | Create a DIAMOND formatted reference database for the specified taxon. |
| -eggnog-diamond-search | method | Run eggNOG search using Diamond aligner. |
| -eggnog-hmmer-search | method | Run eggNOG search using HMMER aligner. |
| -eggnog-feature-table | method | Create an eggnog table. |
| -eggnog-annotate | method | Annotate orthologs against eggNOG database. |
| predict-genes-prodigal | method | Predict gene sequences from MAGs or contigs using Prodigal. |
| fetch-kaiju-db | method | Fetch Kaiju database. |
| -classify-kaiju | method | Classify sequences using Kaiju. |
| fetch-eggnog-hmmer-db | method | Fetch the taxon specific database necessary to run the eggnog-hmmer-search action. |
| extract-annotations | method | Extract annotation frequencies from all annotations. |
| -filter-kraken2-reports-by-abundance | method | Filter kraken2 reports by relative abundance. |
| -filter-kraken2-results-by-metadata | method | Filter Kraken2 reports and outputs. |
| -merge-kraken2-results | method | Merge kraken2 reports and outputs. |
| -align-outputs-with-reports | method | Align unfiltered kraken2 outputs with filtered kraken2 reports. |
| -filter-reads-kraken2 | method | Filter Kraken2-classified reads by taxonomy. |
| -visualize-collapsed-contigs | visualizer | Visualize collapsed contig abundances |
| classify-kraken2 | pipeline | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| map-taxonomy-to-contigs | pipeline | Map contig IDs to taxonomy strings from Kraken 2. |
| collapse-contigs | pipeline | Collapse the contig abundances based on taxonomy. |
| search-orthologs-diamond | pipeline | Run eggNOG search using diamond aligner. |
| search-orthologs-hmmer | pipeline | Run eggNOG search using HMMER aligner. |
| map-eggnog | pipeline | Annotate orthologs against eggNOG database. |
| classify-kaiju | pipeline | Classify sequences using Kaiju. |
| filter-kraken2-results | pipeline | Filter kraken2 reports and outputs by metadata and abundance. |
| filter-reads-kraken2 | pipeline | Filter Kraken2-classified reads by taxonomy. |
Artifact Classes¶
EggnogHmmerIdmap |
Formats¶
EggnogHmmerIdmapFileFmt |
EggnogHmmerIdmapDirectoryFmt |
annotate -classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs]|FeatureData[MAG]|SampleData[MAGs] Sequences to be classified. Single-/paired-end reads, contigs, or assembled MAGs can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate collate-kraken2-reports¶
Collates kraken2 reports.
Inputs¶
- reports:
List[SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_reports:
SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate collate-kraken2-outputs¶
Collates kraken2 outputs.
Inputs¶
- outputs:
List[SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_outputs:
SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate estimate-bracken¶
This method uses Bracken to re-estimate read abundances. Only available on Linux platforms.
Citations¶
Inputs¶
- kraken2_reports:
SampleData[Kraken2Report % Properties('reads')] Reports produced by Kraken2.[required]
- db:
BrackenDB Bracken database.[required]
Parameters¶
- threshold:
Int%Range(0, None) Bracken: number of reads required PRIOR to abundance estimation to perform re-estimation.[default:
0]- read_len:
Int%Range(0, None) Bracken: the ideal length of reads in your sample. For paired end data (e.g., 2x150) this should be set to the length of the single-end reads (e.g., 150).[default:
100]- level:
Str%Choices('D', 'P', 'C', 'O', 'F', 'G', 'S') Bracken: specifies the taxonomic rank to analyze. Each classification at this specified rank will receive an estimated number of reads belonging to that rank after abundance estimation.[default:
'S']- include_unclassified:
Bool Bracken does not include the unclassified read counts in the feature table. Set this to True to include those regardless.[default:
True]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('bracken')] Reports modified by Bracken.[required]
- taxonomy:
FeatureData[Taxonomy] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
annotate build-kraken-db¶
This method builds Kraken 2 and Bracken databases either (1) from provided DNA sequences to build a custom database, or (2) simply fetches pre-built versions from an online resource.
Citations¶
Wood et al., 2019; Lu et al., 2017
Inputs¶
- seqs:
List[FeatureData[Sequence]] Sequences to be added to the Kraken 2 database.[optional]
Parameters¶
- collection:
Str%Choices('viral', 'minusb', 'standard', 'standard8', 'standard16', 'pluspf', 'pluspf8', 'pluspf16', 'pluspfp', 'pluspfp8', 'pluspfp16', 'eupathdb', 'nt', 'corent', 'gtdb', 'greengenes', 'rdp', 'silva132', 'silva138') Name of the database collection to be fetched. Please check https://
benlangmead .github .io /aws -indexes /k2 for the description of the available options.[optional] - threads:
Int%Range(1, None) Number of threads. Only applicable when building a custom database.[default:
1]- kmer_len:
Int%Range(1, None) K-mer length in bp/aa.[default:
35]- minimizer_len:
Int%Range(1, None) Minimizer length in bp/aa.[default:
31]- minimizer_spaces:
Int%Range(1, None) Number of characters in minimizer that are ignored in comparisons.[default:
7]- no_masking:
Bool Avoid masking low-complexity sequences prior to building; masking requires dustmasker or segmasker to be installed in PATH.[default:
False]- max_db_size:
Int%Range(0, None) Maximum number of bytes for Kraken 2 hash table; if the estimator determines more would normally be needed, the reference library will be downsampled to fit.[default:
0]- use_ftp:
Bool Use FTP for downloading instead of RSYNC.[default:
False]- load_factor:
Float%Range(0, 1) Proportion of the hash table to be populated.[default:
0.7]- fast_build:
Bool Do not require database to be deterministically built when using multiple threads. This is faster, but does introduce variability in minimizer/LCA pairs.[default:
False]- read_len:
List[Int%Range(1, None)] Ideal read lengths to be used while building the Bracken database.[optional]
Outputs¶
annotate inspect-kraken2-db¶
This method generates a report of identical format to those generated by classify_kraken2, with a slightly different interpretation. Instead of reporting the number of inputs classified to a taxon/clade, the report displays the number of minimizers mapped to each taxon/clade.
Citations¶
Inputs¶
- db:
Kraken2DB The Kraken 2 database for which to generate the report.[required]
Parameters¶
Outputs¶
- report:
Kraken2DBReport The report of the supplied database.[required]
annotate kraken2-to-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
SampleData[Kraken2Report] Per-sample Kraken 2 reports.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- table:
FeatureTable[PresenceAbsence] A presence/absence table of selected features. The features are not of even ranks, but will be the most specific rank available.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate kraken2-to-mag-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
FeatureData[Kraken2Report % Properties('mags')] Per-sample Kraken 2 reports.[required]
- outputs:
FeatureData[Kraken2Output % Properties('mags')] Per-sample Kraken 2 output files.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate build-custom-diamond-db¶
Creates an artifact containing a binary DIAMOND database file (ref_db.dmnd) from a protein reference database file in FASTA format.
Citations¶
Inputs¶
- seqs:
FeatureData[ProteinSequence] Protein reference database.[required]
- taxonomy:
ReferenceDB[NCBITaxonomy] Reference taxonomy, needed to provide taxonomy features.[optional]
Parameters¶
- threads:
Int%Range(1, None) Number of CPU threads.[default:
1]- file_buffer_size:
Int%Range(1, None) File buffer size in bytes.[default:
67108864]- ignore_warnings:
Bool Ignore warnings.[default:
False]- no_parse_seqids:
Bool Print raw seqids without parsing.[default:
False]
Outputs¶
- db:
ReferenceDB[Diamond] DIAMOND database.[required]
annotate fetch-eggnog-db¶
Downloads EggNOG reference database using the download_eggnog_data.py script from eggNOG. Here, this script downloads 3 files and stores them in the output artifact. At least 80 GB of storage space is required to run this action.
Citations¶
Outputs¶
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
annotate fetch-diamond-db¶
Downloads Diamond reference database. This action downloads 1 file (ref_db.dmnd). At least 18 GB of storage space is required to run this action.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database.[required]
annotate fetch-eggnog-proteins¶
Downloads eggNOG proteome database. This script downloads 2 files (e5.proteomes.faa and e5.taxid_info.tsv) and creates and artifact with them. At least 18 GB of storage space is required to run this action.
Citations¶
Outputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information.[required]
annotate fetch-ncbi-taxonomy¶
Downloads NCBI reference taxonomy from the NCBI FTP server. The resulting artifact is required by the build-custom-diamond-db action if one wishes to create a Diamond data base with taxonomy features. At least 30 GB of storage space is required to run this action.
Citations¶
National Center for Biotechnology Information (NCBI), n.d.
Outputs¶
- taxonomy:
ReferenceDB[NCBITaxonomy] NCBI reference taxonomy.[required]
annotate build-eggnog-diamond-db¶
Creates a DIAMOND database which contains the protein sequences that belong to the specified taxon.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information (generated through the
fetch-eggnog-proteinsaction).[required]
Parameters¶
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database for the specified taxon.[required]
annotate -eggnog-diamond-search¶
This method performs the steps by which we find our possible target sequences to annotate using the Diamond search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for ortholog hits.[required]
- db:
ReferenceDB[Diamond] Diamond database.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-hmmer-search¶
This method performs the steps by which we find our possible target sequences to annotate using the HMMER search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-feature-table¶
Create an eggnog table.
Inputs¶
- seed_orthologs:
SampleData[Orthologs] Sequence data to be turned into an eggnog feature table.[required]
Outputs¶
- table:
FeatureTable[Frequency] <no description>[required]
annotate -eggnog-annotate¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] <no description>[required]
- db:
ReferenceDB[Eggnog] <no description>[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggNOG database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] <no description>[required]
annotate predict-genes-prodigal¶
Prodigal (PROkaryotic DYnamic programming Gene-finding ALgorithm), a gene prediction algorithm designed for improved gene structure prediction, translation initiation site recognition, and reduced false positives in bacterial and archaeal genomes.
Citations¶
Inputs¶
- seqs:
FeatureData[MAG]|SampleData[MAGs]|SampleData[Contigs] MAGs or contigs for which one wishes to predict genes.[required]
Parameters¶
- translation_table_number:
Str%Choices('1', '2', '3', '4', '5', '6', '9', '10', '11', '12', '13', '14', '15', '16', '21', '22', '23', '24', '25') Translation table to be used to translate genes into sequences of amino acids. See https://
www .ncbi .nlm .nih .gov /Taxonomy /Utils /wprintgc .cgi for reference.[default: '11']- mode:
Str%Choices('single', 'meta') Gene prediction mode. 'single' is suitable for single genome analysis (e.g., MAGs), 'meta' is suitable for metagenome analysis (e.g., contigs from mixed communities).[default:
'single']- closed:
Bool Treat sequences as complete genomes with closed ends. Use this for finished genomes where the sequences represent complete chromosomes or plasmids.[default:
False]- no_shine_dalgarno:
Bool Bypass Shine-Dalgarno trainer and use a more generic model. Useful for virus, phage, or plasmid sequences that may not follow standard prokaryotic gene patterns.[default:
False]- mask:
Bool Treat runs of N as masked sequence. Useful for assemblies that contain gap regions represented by stretches of N nucleotides.[default:
False]
Outputs¶
- loci:
GenomeData[Loci] Gene coordinates files (one per MAG or sample) listing the location of each predicted gene as well as some additional scoring information.[required]
- genes:
GenomeData[Genes] Fasta files (one per MAG or sample) with the nucleotide sequences of the predicted genes.[required]
- proteins:
GenomeData[Proteins] Fasta files (one per MAG or sample) with the protein translation of the predicted genes.[required]
annotate fetch-kaiju-db¶
This method fetches the latest Kaiju database from Kaiju's web server.
Citations¶
Parameters¶
- database_type:
Str%Choices('nr', 'nr_euk', 'refseq', 'refseq_ref', 'refseq_nr', 'fungi', 'viruses', 'plasmids', 'progenomes', 'rvdb') Type of database to be downloaded. For more information on available types please see the list on Kaiju's web server: https://
bioinformatics -centre .github .io /kaiju /downloads .html[required]
Outputs¶
- db:
KaijuDB Kaiju database.[required]
annotate -classify-kaiju¶
This method uses Kaiju to perform taxonomic classification of DNA sequence reads or contigs.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate fetch-eggnog-hmmer-db¶
Downloads Profile HMM database for the specified taxon.
Citations¶
Huerta-Cepas et al., 2019; HMMER, 2024
Parameters¶
Outputs¶
- idmap:
EggnogHmmerIdmap%Properties('eggnog') List of protein families in
hmm_db.[required]- hmm_db:
ProfileHMM[MultipleProtein]%Properties('eggnog') Collection of Profile HMMs.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein]%Properties('eggnog') Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins]%Properties('eggnog') Seed alignments for the protein families in
hmm_db.[required]
annotate extract-annotations¶
This method extract a specific annotation from the table generated by EggNOG and calculates its frequencies across all contigs (annotation_counts_per_contig) and MAGs (annotation_counts_per_genome).
Inputs¶
- ortholog_annotations:
GenomeData[NOG] Ortholog annotations.[required]
Parameters¶
- annotation:
Str%Choices('cog', 'caz', 'kegg_ko', 'kegg_pathway', 'kegg_reaction', 'kegg_module', 'brite', 'ec') Annotation to extract.[required]
- max_evalue:
Float%Range(0, None) <no description>[default:
1.0]- min_score:
Float%Range(0, None) <no description>[default:
0.0]
Outputs¶
- annotation_counts_per_genome:
FeatureTable[Frequency] Feature table with frequency of each annotation per genome.[required]
- annotation_map:
FeatureMap[FunctionToContigs] Feature map with function to contigs mapping.[required]
- annotation_counts_per_contig:
FeatureTable[Frequency] Feature table with frequency of each annotation per contig.[required]
annotate -filter-kraken2-reports-by-abundance¶
Filters kraken2 reports on a per-taxon basis by relative abundance (relative frequency). Useful for removing suspected spurious classifications.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter by relative abundance.[required]
Parameters¶
- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report.[required]
- remove_empty:
Bool If True, reports with only unclassified reads remaining will be removed from the filtered data.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The relative abundance-filtered kraken2 reports[required]
annotate -filter-kraken2-results-by-metadata¶
Filter Kraken2 reports and outputs based on metadata or remove reports with 100% unclassified reads.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The Kraken reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The Kraken outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] <no description>[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] <no description>[required]
annotate -merge-kraken2-results¶
Merge multiple kraken2 reports and outputs such that the results contain a union of the samples represented in the inputs. If sample IDs overlap across the inputs, these reports and outputs will be processed into a single report or output per sample ID.
Inputs¶
- reports:
List[SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')]] The kraken2 reports to merge. Only reports with the same sample ID are merged into one report.[required]
- outputs:
List[SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')]] The kraken2 outputs to merge. Only outputs with the same sample ID are merged into one output.[required]
Outputs¶
- merged_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The merged kraken2 reports.[required]
- merged_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The merged kraken2 outputs.[required]
annotate -align-outputs-with-reports¶
Inputs¶
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to align with the filtered reports.[required]
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
Outputs¶
- aligned_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The report-aligned filtered kraken2 outputs.[required]
annotate -filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
annotate -visualize-collapsed-contigs¶
Generates a visualization showing histograms of contig abundance distributions per taxon per sample from the original (pre-collapsed) table.
Inputs¶
- table:
FeatureTable[Frequency] Original feature table with contig IDs as feature IDs (before collapsing).[required]
- collapsed_table:
FeatureTable[Frequency] Feature table with contig IDs collapsed to taxonomy IDs.[required]
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings instead of IDs.[optional]
Outputs¶
- visualization:
Visualization <no description>[required]
annotate classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
List[SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]]|List[SampleData[Contigs]]|List[FeatureData[MAG]]|List[SampleData[MAGs]] Sequences to be classified. Single-/paired-end reads,contigs, or assembled MAGs, can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate map-taxonomy-to-contigs¶
Maps contig IDs to their full taxonomy strings based on Kraken 2 classifications. This action processes Kraken 2 reports and outputs produced from contig sequences to create a taxonomy mapping where each contig ID is associated with its assigned taxonomy.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('contigs')] Per-sample Kraken 2 reports for contigs.[required]
- outputs:
SampleData[Kraken2Output % Properties('contigs')] Per-sample Kraken 2 output files for contigs.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- feature_map:
FeatureMap[TaxonomyToContigs] Taxonomy assignments for contigs. Each contig ID is mapped to its full taxonomy string based on Kraken2 classifications. Unclassified contigs are assigned 'd__Unclassified'. Assignments below the provided coverage threshold will be treated as 'd__Unclassified'.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate collapse-contigs¶
This action collapses contig abundances based on their taxonomy assignments. Contig abundances for contigs with the same taxonomy assignment will be averaged.
Inputs¶
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- table:
FeatureTable[Frequency] Table of contig abundances per sample. Feature IDs will be replaced with taxonomy strings.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings in the visualization instead of taxonomy IDs.[optional]
Outputs¶
- collapsed_table:
FeatureTable[Frequency] Contig abundance table with contig IDs replaced by their taxonomy strings. Contigs with the same taxonomy assignment will have their abundances averaged.[required]
- visualization:
Visualization Interactive visualization showing histograms of contig abundance distributions per taxon per sample.[required]
annotate search-orthologs-diamond¶
Use Diamond and eggNOG to align contig or MAG sequences against the Diamond database.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits using the Diamond Database[required]
- db:
ReferenceDB[Diamond] The filepath to an artifact containing the Diamond database[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate search-orthologs-hmmer¶
This method uses HMMER to find possible target sequences to annotate with eggNOG-mapper.
Citations¶
HMMER, 2024; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of profile HMMs in binary format and indexed.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate map-eggnog¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] BLAST6-like table(s) describing the identified orthologs.[required]
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggnog database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] Annotated hits.[required]
annotate classify-kaiju¶
This method uses Kaiju to perform taxonomic classification.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate filter-kraken2-results¶
Filter kraken2 reports and outputs by sample metadata, and/or filter classified taxa by relative abundance.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report. If a taxon is filtered from the report, its associated read counts are removed entirely from the report (i.e., the subtraction of those counts is propagated to parent taxonomic groupings).[optional]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The filtered kraken2 outputs.[required]
annotate filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[default:
1]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
MOdular SHotgun metagenome Pipelines with Integrated provenance Tracking: QIIME 2 plugin gor metagenome analysis withtools for genome binning and functional annotation.
- version:
2026.7.0 - website: https://
github .com /bokulich -lab /q2 -annotate - user support:
- Please post to the QIIME 2 forum for help with this plugin: https://
forum .qiime2 .org
Actions¶
| Name | Type | Short Description |
|---|---|---|
| -classify-kraken2 | method | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| collate-kraken2-reports | method | Collate kraken2 reports. |
| collate-kraken2-outputs | method | Collate kraken2 outputs. |
| estimate-bracken | method | Perform read abundance re-estimation using Bracken. |
| build-kraken-db | method | Build Kraken 2 database. |
| inspect-kraken2-db | method | Inspect a Kraken 2 database. |
| kraken2-to-features | method | Select downstream features from Kraken 2. |
| kraken2-to-mag-features | method | Select downstream MAG features from Kraken 2. |
| build-custom-diamond-db | method | Create a DIAMOND formatted reference database from a FASTA input file. |
| fetch-eggnog-db | method | Fetch the databases necessary to run the eggnog-annotate action. |
| fetch-diamond-db | method | Fetch the complete Diamond database necessary to run the eggnog-diamond-search action. |
| fetch-eggnog-proteins | method | Fetch the databases necessary to run the build-eggnog-diamond-db action. |
| fetch-ncbi-taxonomy | method | Fetch NCBI reference taxonomy. |
| build-eggnog-diamond-db | method | Create a DIAMOND formatted reference database for the specified taxon. |
| -eggnog-diamond-search | method | Run eggNOG search using Diamond aligner. |
| -eggnog-hmmer-search | method | Run eggNOG search using HMMER aligner. |
| -eggnog-feature-table | method | Create an eggnog table. |
| -eggnog-annotate | method | Annotate orthologs against eggNOG database. |
| predict-genes-prodigal | method | Predict gene sequences from MAGs or contigs using Prodigal. |
| fetch-kaiju-db | method | Fetch Kaiju database. |
| -classify-kaiju | method | Classify sequences using Kaiju. |
| fetch-eggnog-hmmer-db | method | Fetch the taxon specific database necessary to run the eggnog-hmmer-search action. |
| extract-annotations | method | Extract annotation frequencies from all annotations. |
| -filter-kraken2-reports-by-abundance | method | Filter kraken2 reports by relative abundance. |
| -filter-kraken2-results-by-metadata | method | Filter Kraken2 reports and outputs. |
| -merge-kraken2-results | method | Merge kraken2 reports and outputs. |
| -align-outputs-with-reports | method | Align unfiltered kraken2 outputs with filtered kraken2 reports. |
| -filter-reads-kraken2 | method | Filter Kraken2-classified reads by taxonomy. |
| -visualize-collapsed-contigs | visualizer | Visualize collapsed contig abundances |
| classify-kraken2 | pipeline | Perform taxonomic classification of reads or MAGs using Kraken 2. |
| map-taxonomy-to-contigs | pipeline | Map contig IDs to taxonomy strings from Kraken 2. |
| collapse-contigs | pipeline | Collapse the contig abundances based on taxonomy. |
| search-orthologs-diamond | pipeline | Run eggNOG search using diamond aligner. |
| search-orthologs-hmmer | pipeline | Run eggNOG search using HMMER aligner. |
| map-eggnog | pipeline | Annotate orthologs against eggNOG database. |
| classify-kaiju | pipeline | Classify sequences using Kaiju. |
| filter-kraken2-results | pipeline | Filter kraken2 reports and outputs by metadata and abundance. |
| filter-reads-kraken2 | pipeline | Filter Kraken2-classified reads by taxonomy. |
Artifact Classes¶
EggnogHmmerIdmap |
Formats¶
EggnogHmmerIdmapFileFmt |
EggnogHmmerIdmapDirectoryFmt |
annotate -classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs]|FeatureData[MAG]|SampleData[MAGs] Sequences to be classified. Single-/paired-end reads, contigs, or assembled MAGs can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate collate-kraken2-reports¶
Collates kraken2 reports.
Inputs¶
- reports:
List[SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_reports:
SampleData[Kraken2Report % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate collate-kraken2-outputs¶
Collates kraken2 outputs.
Inputs¶
- outputs:
List[SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)]] <no description>[required]
Outputs¶
- collated_outputs:
SampleData[Kraken2Output % (Properties('reads', 'contigs', 'mags')¹ | Properties('reads', 'contigs')² | Properties('reads', 'mags')³ | Properties('contigs', 'mags')⁴ | Properties('reads')⁵ | Properties('contigs')⁶ | Properties('mags')⁷)] <no description>[required]
annotate estimate-bracken¶
This method uses Bracken to re-estimate read abundances. Only available on Linux platforms.
Citations¶
Inputs¶
- kraken2_reports:
SampleData[Kraken2Report % Properties('reads')] Reports produced by Kraken2.[required]
- db:
BrackenDB Bracken database.[required]
Parameters¶
- threshold:
Int%Range(0, None) Bracken: number of reads required PRIOR to abundance estimation to perform re-estimation.[default:
0]- read_len:
Int%Range(0, None) Bracken: the ideal length of reads in your sample. For paired end data (e.g., 2x150) this should be set to the length of the single-end reads (e.g., 150).[default:
100]- level:
Str%Choices('D', 'P', 'C', 'O', 'F', 'G', 'S') Bracken: specifies the taxonomic rank to analyze. Each classification at this specified rank will receive an estimated number of reads belonging to that rank after abundance estimation.[default:
'S']- include_unclassified:
Bool Bracken does not include the unclassified read counts in the feature table. Set this to True to include those regardless.[default:
True]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('bracken')] Reports modified by Bracken.[required]
- taxonomy:
FeatureData[Taxonomy] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
annotate build-kraken-db¶
This method builds Kraken 2 and Bracken databases either (1) from provided DNA sequences to build a custom database, or (2) simply fetches pre-built versions from an online resource.
Citations¶
Wood et al., 2019; Lu et al., 2017
Inputs¶
- seqs:
List[FeatureData[Sequence]] Sequences to be added to the Kraken 2 database.[optional]
Parameters¶
- collection:
Str%Choices('viral', 'minusb', 'standard', 'standard8', 'standard16', 'pluspf', 'pluspf8', 'pluspf16', 'pluspfp', 'pluspfp8', 'pluspfp16', 'eupathdb', 'nt', 'corent', 'gtdb', 'greengenes', 'rdp', 'silva132', 'silva138') Name of the database collection to be fetched. Please check https://
benlangmead .github .io /aws -indexes /k2 for the description of the available options.[optional] - threads:
Int%Range(1, None) Number of threads. Only applicable when building a custom database.[default:
1]- kmer_len:
Int%Range(1, None) K-mer length in bp/aa.[default:
35]- minimizer_len:
Int%Range(1, None) Minimizer length in bp/aa.[default:
31]- minimizer_spaces:
Int%Range(1, None) Number of characters in minimizer that are ignored in comparisons.[default:
7]- no_masking:
Bool Avoid masking low-complexity sequences prior to building; masking requires dustmasker or segmasker to be installed in PATH.[default:
False]- max_db_size:
Int%Range(0, None) Maximum number of bytes for Kraken 2 hash table; if the estimator determines more would normally be needed, the reference library will be downsampled to fit.[default:
0]- use_ftp:
Bool Use FTP for downloading instead of RSYNC.[default:
False]- load_factor:
Float%Range(0, 1) Proportion of the hash table to be populated.[default:
0.7]- fast_build:
Bool Do not require database to be deterministically built when using multiple threads. This is faster, but does introduce variability in minimizer/LCA pairs.[default:
False]- read_len:
List[Int%Range(1, None)] Ideal read lengths to be used while building the Bracken database.[optional]
Outputs¶
annotate inspect-kraken2-db¶
This method generates a report of identical format to those generated by classify_kraken2, with a slightly different interpretation. Instead of reporting the number of inputs classified to a taxon/clade, the report displays the number of minimizers mapped to each taxon/clade.
Citations¶
Inputs¶
- db:
Kraken2DB The Kraken 2 database for which to generate the report.[required]
Parameters¶
Outputs¶
- report:
Kraken2DBReport The report of the supplied database.[required]
annotate kraken2-to-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
SampleData[Kraken2Report] Per-sample Kraken 2 reports.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- table:
FeatureTable[PresenceAbsence] A presence/absence table of selected features. The features are not of even ranks, but will be the most specific rank available.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate kraken2-to-mag-features¶
Convert a Kraken 2 report, which is an annotated NCBI taxonomy tree, into generic artifacts for downstream analyses.
Inputs¶
- reports:
FeatureData[Kraken2Report % Properties('mags')] Per-sample Kraken 2 reports.[required]
- outputs:
FeatureData[Kraken2Output % Properties('mags')] Per-sample Kraken 2 output files.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate build-custom-diamond-db¶
Creates an artifact containing a binary DIAMOND database file (ref_db.dmnd) from a protein reference database file in FASTA format.
Citations¶
Inputs¶
- seqs:
FeatureData[ProteinSequence] Protein reference database.[required]
- taxonomy:
ReferenceDB[NCBITaxonomy] Reference taxonomy, needed to provide taxonomy features.[optional]
Parameters¶
- threads:
Int%Range(1, None) Number of CPU threads.[default:
1]- file_buffer_size:
Int%Range(1, None) File buffer size in bytes.[default:
67108864]- ignore_warnings:
Bool Ignore warnings.[default:
False]- no_parse_seqids:
Bool Print raw seqids without parsing.[default:
False]
Outputs¶
- db:
ReferenceDB[Diamond] DIAMOND database.[required]
annotate fetch-eggnog-db¶
Downloads EggNOG reference database using the download_eggnog_data.py script from eggNOG. Here, this script downloads 3 files and stores them in the output artifact. At least 80 GB of storage space is required to run this action.
Citations¶
Outputs¶
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
annotate fetch-diamond-db¶
Downloads Diamond reference database. This action downloads 1 file (ref_db.dmnd). At least 18 GB of storage space is required to run this action.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database.[required]
annotate fetch-eggnog-proteins¶
Downloads eggNOG proteome database. This script downloads 2 files (e5.proteomes.faa and e5.taxid_info.tsv) and creates and artifact with them. At least 18 GB of storage space is required to run this action.
Citations¶
Outputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information.[required]
annotate fetch-ncbi-taxonomy¶
Downloads NCBI reference taxonomy from the NCBI FTP server. The resulting artifact is required by the build-custom-diamond-db action if one wishes to create a Diamond data base with taxonomy features. At least 30 GB of storage space is required to run this action.
Citations¶
National Center for Biotechnology Information (NCBI), n.d.
Outputs¶
- taxonomy:
ReferenceDB[NCBITaxonomy] NCBI reference taxonomy.[required]
annotate build-eggnog-diamond-db¶
Creates a DIAMOND database which contains the protein sequences that belong to the specified taxon.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- eggnog_proteins:
ReferenceDB[EggnogProteinSequences] eggNOG database of protein sequences and their corresponding taxonomy information (generated through the
fetch-eggnog-proteinsaction).[required]
Parameters¶
Outputs¶
- db:
ReferenceDB[Diamond] Complete Diamond reference database for the specified taxon.[required]
annotate -eggnog-diamond-search¶
This method performs the steps by which we find our possible target sequences to annotate using the Diamond search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for ortholog hits.[required]
- db:
ReferenceDB[Diamond] Diamond database.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-hmmer-search¶
This method performs the steps by which we find our possible target sequences to annotate using the HMMER search functionality from the eggnog emapper.py script.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] BLAST6-like table(s) describing the identified orthologs. One table per sample or MAG in the input.[required]
- table:
FeatureTable[Frequency] Feature table with counts of orthologs per sample/MAG.[required]
- loci:
GenomeData[Loci] Loci of the identified orthologs.[required]
annotate -eggnog-feature-table¶
Create an eggnog table.
Inputs¶
- seed_orthologs:
SampleData[Orthologs] Sequence data to be turned into an eggnog feature table.[required]
Outputs¶
- table:
FeatureTable[Frequency] <no description>[required]
annotate -eggnog-annotate¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] <no description>[required]
- db:
ReferenceDB[Eggnog] <no description>[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggNOG database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] <no description>[required]
annotate predict-genes-prodigal¶
Prodigal (PROkaryotic DYnamic programming Gene-finding ALgorithm), a gene prediction algorithm designed for improved gene structure prediction, translation initiation site recognition, and reduced false positives in bacterial and archaeal genomes.
Citations¶
Inputs¶
- seqs:
FeatureData[MAG]|SampleData[MAGs]|SampleData[Contigs] MAGs or contigs for which one wishes to predict genes.[required]
Parameters¶
- translation_table_number:
Str%Choices('1', '2', '3', '4', '5', '6', '9', '10', '11', '12', '13', '14', '15', '16', '21', '22', '23', '24', '25') Translation table to be used to translate genes into sequences of amino acids. See https://
www .ncbi .nlm .nih .gov /Taxonomy /Utils /wprintgc .cgi for reference.[default: '11']- mode:
Str%Choices('single', 'meta') Gene prediction mode. 'single' is suitable for single genome analysis (e.g., MAGs), 'meta' is suitable for metagenome analysis (e.g., contigs from mixed communities).[default:
'single']- closed:
Bool Treat sequences as complete genomes with closed ends. Use this for finished genomes where the sequences represent complete chromosomes or plasmids.[default:
False]- no_shine_dalgarno:
Bool Bypass Shine-Dalgarno trainer and use a more generic model. Useful for virus, phage, or plasmid sequences that may not follow standard prokaryotic gene patterns.[default:
False]- mask:
Bool Treat runs of N as masked sequence. Useful for assemblies that contain gap regions represented by stretches of N nucleotides.[default:
False]
Outputs¶
- loci:
GenomeData[Loci] Gene coordinates files (one per MAG or sample) listing the location of each predicted gene as well as some additional scoring information.[required]
- genes:
GenomeData[Genes] Fasta files (one per MAG or sample) with the nucleotide sequences of the predicted genes.[required]
- proteins:
GenomeData[Proteins] Fasta files (one per MAG or sample) with the protein translation of the predicted genes.[required]
annotate fetch-kaiju-db¶
This method fetches the latest Kaiju database from Kaiju's web server.
Citations¶
Parameters¶
- database_type:
Str%Choices('nr', 'nr_euk', 'refseq', 'refseq_ref', 'refseq_nr', 'fungi', 'viruses', 'plasmids', 'progenomes', 'rvdb') Type of database to be downloaded. For more information on available types please see the list on Kaiju's web server: https://
bioinformatics -centre .github .io /kaiju /downloads .html[required]
Outputs¶
- db:
KaijuDB Kaiju database.[required]
annotate -classify-kaiju¶
This method uses Kaiju to perform taxonomic classification of DNA sequence reads or contigs.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate fetch-eggnog-hmmer-db¶
Downloads Profile HMM database for the specified taxon.
Citations¶
Huerta-Cepas et al., 2019; HMMER, 2024
Parameters¶
Outputs¶
- idmap:
EggnogHmmerIdmap%Properties('eggnog') List of protein families in
hmm_db.[required]- hmm_db:
ProfileHMM[MultipleProtein]%Properties('eggnog') Collection of Profile HMMs.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein]%Properties('eggnog') Collection of Profile HMMs in binary format and indexed.[required]
- seed_alignments:
GenomeData[Proteins]%Properties('eggnog') Seed alignments for the protein families in
hmm_db.[required]
annotate extract-annotations¶
This method extract a specific annotation from the table generated by EggNOG and calculates its frequencies across all contigs (annotation_counts_per_contig) and MAGs (annotation_counts_per_genome).
Inputs¶
- ortholog_annotations:
GenomeData[NOG] Ortholog annotations.[required]
Parameters¶
- annotation:
Str%Choices('cog', 'caz', 'kegg_ko', 'kegg_pathway', 'kegg_reaction', 'kegg_module', 'brite', 'ec') Annotation to extract.[required]
- max_evalue:
Float%Range(0, None) <no description>[default:
1.0]- min_score:
Float%Range(0, None) <no description>[default:
0.0]
Outputs¶
- annotation_counts_per_genome:
FeatureTable[Frequency] Feature table with frequency of each annotation per genome.[required]
- annotation_map:
FeatureMap[FunctionToContigs] Feature map with function to contigs mapping.[required]
- annotation_counts_per_contig:
FeatureTable[Frequency] Feature table with frequency of each annotation per contig.[required]
annotate -filter-kraken2-reports-by-abundance¶
Filters kraken2 reports on a per-taxon basis by relative abundance (relative frequency). Useful for removing suspected spurious classifications.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter by relative abundance.[required]
Parameters¶
- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report.[required]
- remove_empty:
Bool If True, reports with only unclassified reads remaining will be removed from the filtered data.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The relative abundance-filtered kraken2 reports[required]
annotate -filter-kraken2-results-by-metadata¶
Filter Kraken2 reports and outputs based on metadata or remove reports with 100% unclassified reads.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The Kraken reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The Kraken outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] <no description>[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] <no description>[required]
annotate -merge-kraken2-results¶
Merge multiple kraken2 reports and outputs such that the results contain a union of the samples represented in the inputs. If sample IDs overlap across the inputs, these reports and outputs will be processed into a single report or output per sample ID.
Inputs¶
- reports:
List[SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')]] The kraken2 reports to merge. Only reports with the same sample ID are merged into one report.[required]
- outputs:
List[SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')]] The kraken2 outputs to merge. Only outputs with the same sample ID are merged into one output.[required]
Outputs¶
- merged_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The merged kraken2 reports.[required]
- merged_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The merged kraken2 outputs.[required]
annotate -align-outputs-with-reports¶
Inputs¶
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to align with the filtered reports.[required]
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
Outputs¶
- aligned_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The report-aligned filtered kraken2 outputs.[required]
annotate -filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
annotate -visualize-collapsed-contigs¶
Generates a visualization showing histograms of contig abundance distributions per taxon per sample from the original (pre-collapsed) table.
Inputs¶
- table:
FeatureTable[Frequency] Original feature table with contig IDs as feature IDs (before collapsing).[required]
- collapsed_table:
FeatureTable[Frequency] Feature table with contig IDs collapsed to taxonomy IDs.[required]
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings instead of IDs.[optional]
Outputs¶
- visualization:
Visualization <no description>[required]
annotate classify-kraken2¶
Use Kraken 2 to classify provided DNA sequence reads, contigs, or MAGs into taxonomic groups.
Citations¶
Inputs¶
- seqs:
List[SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]]|List[SampleData[Contigs]]|List[FeatureData[MAG]]|List[SampleData[MAGs]] Sequences to be classified. Single-/paired-end reads,contigs, or assembled MAGs, can be provided.[required]
- db:
Kraken2DB Kraken 2 database.[required]
Parameters¶
- threads:
Int%Range(1, None) Number of threads.[default:
1]- confidence:
Float%Range(0, 1, inclusive_end=True) Confidence score threshold.[default:
0.0]- minimum_base_quality:
Int%Range(0, None) Minimum base quality used in classification. Only applies when reads are used as input.[default:
0]- memory_mapping:
Bool Avoids loading the database into RAM.[default:
False]- minimum_hit_groups:
Int%Range(1, None) Minimum number of hit groups (overlapping k-mers sharing the same minimizer).[default:
2]- quick:
Bool Quick operation (use first hit or hits).[default:
False]- report_minimizer_data:
Bool Include number of read-minimizers per-taxon and unique read-minimizers per-taxon in the report. If this parameter is enabled then merging kraken2 reports with the same sample ID from two or more input artifacts will not be possible.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- reports:
SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|FeatureData[Kraken2Report % Properties('mags')]|SampleData[Kraken2Report % Properties('mags')] Reports produced by Kraken2.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|FeatureData[Kraken2Output % Properties('mags')]|SampleData[Kraken2Output % Properties('mags')] Output files produced by Kraken2.[required]
annotate map-taxonomy-to-contigs¶
Maps contig IDs to their full taxonomy strings based on Kraken 2 classifications. This action processes Kraken 2 reports and outputs produced from contig sequences to create a taxonomy mapping where each contig ID is associated with its assigned taxonomy.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('contigs')] Per-sample Kraken 2 reports for contigs.[required]
- outputs:
SampleData[Kraken2Output % Properties('contigs')] Per-sample Kraken 2 output files for contigs.[required]
Parameters¶
- coverage_threshold:
Float%Range(0, 100, inclusive_end=True) The minimum percent coverage required to produce a feature.[default:
0.1]
Outputs¶
- feature_map:
FeatureMap[TaxonomyToContigs] Taxonomy assignments for contigs. Each contig ID is mapped to its full taxonomy string based on Kraken2 classifications. Unclassified contigs are assigned 'd__Unclassified'. Assignments below the provided coverage threshold will be treated as 'd__Unclassified'.[required]
- taxonomy:
FeatureData[Taxonomy] Output taxonomy. Infra-clade ranks are ignored unless if they are strain-level. Missing internal ranks are annotated by their next most specific rank, with the exception of k__Bacteria and k__Archaea, which match their domain name.[required]
annotate collapse-contigs¶
This action collapses contig abundances based on their taxonomy assignments. Contig abundances for contigs with the same taxonomy assignment will be averaged.
Inputs¶
- contig_map:
FeatureMap[TaxonomyToContigs] Mapping between contig IDs and assigned taxonomy IDs.[required]
- table:
FeatureTable[Frequency] Table of contig abundances per sample. Feature IDs will be replaced with taxonomy strings.[required]
- taxonomy:
FeatureData[Taxonomy] Optional taxonomy mapping to display taxonomy strings in the visualization instead of taxonomy IDs.[optional]
Outputs¶
- collapsed_table:
FeatureTable[Frequency] Contig abundance table with contig IDs replaced by their taxonomy strings. Contigs with the same taxonomy assignment will have their abundances averaged.[required]
- visualization:
Visualization Interactive visualization showing histograms of contig abundance distributions per taxon per sample.[required]
annotate search-orthologs-diamond¶
Use Diamond and eggNOG to align contig or MAG sequences against the Diamond database.
Citations¶
Buchfink et al., 2021; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits using the Diamond Database[required]
- db:
ReferenceDB[Diamond] The filepath to an artifact containing the Diamond database[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate search-orthologs-hmmer¶
This method uses HMMER to find possible target sequences to annotate with eggNOG-mapper.
Citations¶
HMMER, 2024; Huerta-Cepas et al., 2019
Inputs¶
- seqs:
SampleData[Contigs]|SampleData[MAGs]|FeatureData[MAG] Sequences to be searched for hits.[required]
- pressed_hmm_db:
ProfileHMM[PressedProtein] Collection of profile HMMs in binary format and indexed.[required]
- idmap:
EggnogHmmerIdmap List of protein families in
pressed_hmm_db.[required]- seed_alignments:
GenomeData[Proteins] Seed alignments for the protein families in
pressed_hmm_db.[required]
Parameters¶
- num_cpus:
Int Number of CPUs to utilize per partition. '0' will use all available.[default:
1]- db_in_memory:
Bool Read database into memory. The database can be very large, so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- eggnog_hits:
SampleData[Orthologs % (Properties('contigs')¹ | Properties('mags')²'³)] <no description>[required]
- table:
FeatureTable[Frequency] <no description>[required]
- loci:
GenomeData[Loci] <no description>[required]
annotate map-eggnog¶
Apply eggnog mapper to annotate seed orthologs.
Citations¶
Inputs¶
- eggnog_hits:
SampleData[Orthologs % Properties('contigs', 'mags')¹ | Orthologs % Properties('contigs')² | Orthologs % Properties('mags')³ | Orthologs⁴] BLAST6-like table(s) describing the identified orthologs.[required]
- db:
ReferenceDB[Eggnog] eggNOG annotation database.[required]
Parameters¶
- db_in_memory:
Bool Read eggnog database into memory. The eggnog database is very large (>44GB), so this option should only be used on clusters or other machines with enough memory.[default:
False]- num_cpus:
Int%Range(0, None) Number of CPUs to utilize. '0' will use all available.[default:
1]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- ortholog_annotations:
GenomeData[NOG % Properties('contigs', 'mags')¹ | NOG % Properties('contigs')² | NOG % Properties('mags')³ | NOG⁴] Annotated hits.[required]
annotate classify-kaiju¶
This method uses Kaiju to perform taxonomic classification.
Citations¶
Inputs¶
- seqs:
SampleData[SequencesWithQuality | PairedEndSequencesWithQuality | JoinedSequencesWithQuality]|SampleData[Contigs] Sequences to be classified.[required]
- db:
KaijuDB Kaiju database.[required]
Parameters¶
- z:
Int%Range(1, None) Number of threads.[default:
1]- a:
Str%Choices('greedy', 'mem') Run mode.[default:
'greedy']- e:
Int%Range(1, None) Number of mismatches allowed in Greedy mode.[default:
3]- m:
Int%Range(1, None) Minimum match length.[default:
11]- s:
Int%Range(1, None) Minimum match score in Greedy mode.[default:
65]- evalue:
Float%Range(0, 1) Minimum E-value in Greedy mode.[default:
0.01]- x:
Bool Enable SEG low complexity filter.[default:
True]- r:
Str%Choices('phylum', 'class', 'order', 'family', 'genus', 'species') Taxonomic rank.[default:
'species']- c:
Float%Range(0, 100) Minimum required number or fraction of reads for the taxon (except viruses) to be reported.[default:
0.0]- exp:
Bool Expand viruses, which are always shown as full taxon path and read counts are not summarized in higher taxonomic levels.[default:
False]- u:
Bool Do not count unclassified reads for the total reads when calculating percentages for classified reads.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[optional]
Outputs¶
- abundances:
FeatureTable[Frequency] Sequence abundances.[required]
- taxonomy:
FeatureData[Taxonomy] Linked taxonomy.[required]
annotate filter-kraken2-results¶
Filter kraken2 reports and outputs by sample metadata, and/or filter classified taxa by relative abundance.
Inputs¶
- reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The kraken2 reports to filter.[required]
- outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The kraken2 outputs to filter.[required]
Parameters¶
- metadata:
Metadata Metadata indicating which IDs to filter. The optional
whereparameter may be used to filter IDs based on specified conditions in the metadata. The optionalexclude_idsparameter may be used to exclude the IDs specified in the metadata from the filter.[optional]- where:
Str Optional SQLite WHERE clause specifying metadata criteria that must be met to be included in the filtered data. If not provided, all IDs in
metadatathat are also in the data will be retained.[optional]- exclude_ids:
Bool If True, the samples selected by the
metadataand optionalwhereparameter will be excluded from the filtered data.[default:False]- remove_empty:
Bool If True, reports with only unclassified reads will be removed from the filtered data. Reports containing sequences classified only as root will also be removed.[default:
False]- abundance_threshold:
Float%Range(0, 1, inclusive_end=True) A proportion between 0 and 1 representing the minimum relative abundance (by classified read count) that a taxon must have to be retained in the filtered report. If a taxon is filtered from the report, its associated read counts are removed entirely from the report (i.e., the subtraction of those counts is propagated to parent taxonomic groupings).[optional]
Outputs¶
- filtered_reports:
SampleData[Kraken2Report % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Report % Properties('contigs', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'mags')]|SampleData[Kraken2Report % Properties('reads', 'contigs')]|SampleData[Kraken2Report % Properties('reads')]|SampleData[Kraken2Report % Properties('contigs')]|SampleData[Kraken2Report % Properties('mags')]|FeatureData[Kraken2Report % Properties('mags')] The filtered kraken2 reports.[required]
- filtered_outputs:
SampleData[Kraken2Output % Properties('reads', 'contigs', 'mags')]|SampleData[Kraken2Output % Properties('contigs', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'mags')]|SampleData[Kraken2Output % Properties('reads', 'contigs')]|SampleData[Kraken2Output % Properties('reads')]|SampleData[Kraken2Output % Properties('contigs')]|SampleData[Kraken2Output % Properties('mags')]|FeatureData[Kraken2Output % Properties('mags')] The filtered kraken2 outputs.[required]
annotate filter-reads-kraken2¶
Filter single-end or paired-end reads by matching Kraken2-assigned taxonomy, with optional descendant expansion and inverse filtering.
Inputs¶
- reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] The original reads that were classified by Kraken2. The sample IDs and read headers must match those used to generate
reportsandoutputs.[required]- reports:
SampleData[Kraken2Report % Properties('reads')] Kraken2 reports generated from
reads. Used to identify matching taxa and (optionally) all descendant taxa.[required]- outputs:
SampleData[Kraken2Output % Properties('reads')] Kraken2 per-read outputs generated from
reads. Used to map matched taxa to read IDs that will be filtered.[required]
Parameters¶
- taxonomy:
Str Taxonomy query used for read filtering. Can be a Kraken2 taxon name (for example, "Bacteria") or a taxon ID (for example, "2").[required]
- include_descendants:
Bool If True, include all descendant taxa of each matching taxon.[default:
True]- contains:
Bool If True, match taxon names using case-insensitive substring matching. If False, use exact case-insensitive matching.[default:
False]- exclude:
Bool If False, retain reads that match the taxonomy query. If True, discard matching reads and retain the rest.[default:
False]- num_partitions:
Int%Range(1, None) The number of partitions to split the contigs into. Defaults to partitioning into individual samples.[default:
1]
Outputs¶
- filtered_reads:
SampleData[SequencesWithQuality]|SampleData[PairedEndSequencesWithQuality] Reads filtered according to Kraken2 taxonomy matches.[required]
- Links
- Documentation
- Source Code
- Stars
- 4
- Last Commit
- d4faae6
- Available Distros
- 2026.7
- 2026.7/moshpit
- 2026.4
- 2026.4/moshpit
- 2026.1
- 2026.1/moshpit
- 2025.10
- 2025.10/moshpit
- 2025.7
- 2025.7/moshpit
- 2025.4
- 2025.4/moshpit