Orthologs.Align.MultipleSequenceAlignment
The MultipleSequenceAlignment (MSA) class uses the standard configuration
Usage
Orthologs.Align.MultipleSequenceAlignment(
project=None, project_path=os.getcwd(), genbank=GenBank, **kwargs
)along with function dispatching to give the end-user access to multiple alignment tools.
Methods
| Name | Description |
|---|---|
| __init__() | Initialize the MultipleSequenceAlignment class. |
| clustalo() | Align protein/amino acid sequences using Clustal Omega. |
| guidance2() | Run the GUIDANCE2 command line wrapper from BioPython. |
| pal2nal() | This Pal2Nal method works with the Pal2Nal command line wrapper. |
__init__()
Initialize the MultipleSequenceAlignment class.
Usage
__init__(project=None, project_path=os.getcwd(), genbank=GenBank, **kwargs)Parameters
project: str | None = None-
The project name.
project_path: str | Path = os.getcwd()-
The path to the project.
genbank: class = GenBank-
The composer parameter which is used to configure the GenBank class with the MSA class.
kwargs: dict = {}-
The kwargs are used with the dispatcher as a way to control the alignment pipeline. Can include:
- Guidance_config: Configuration for GUIDANCE2
- Pal2Nal_config: Configuration for PAL2NAL
- ClustalO_config: Configuration for ClustalO
Returns
- If the kwargs are utilized with YAML or other configurations, then this class returns an alignment dictionary, which can be parsed to run specific alignment algorithms.
clustalo()
Align protein/amino acid sequences using Clustal Omega.
Usage
clustalo(infile, outfile, outfmt="fasta")Parameters
infile-
Input a multifasta protein file.
outfile-
Output an aligned multifasta file.
outfmt="fasta"- (Default value = “fasta”)
guidance2()
Run the GUIDANCE2 command line wrapper from BioPython.
Usage
guidance2(
seqFile,
msaProgram,
seqType,
dataset="MSA",
seqFilter=None,
columnFilter=None,
maskFilter=None,
**kwargs
)The Guidance2 algorithm is used to filter sequence alignments in different ways. Here we employ a few of our own strategies on top of Guidance2.
Parameters
seqFile: str | Path-
The sequence file required by GUIDANCE2.
msaProgram: str-
The msa program to be used by GUIDANCE2. (“CLUSTALW”, “PRANK”, “MAFFT”, or “MUSCLE”)
seqType: str-
The type of sequences to be aligned in GUIDANCE2. (“aa”, “nuc”, or “codon”)
dataset: str = "MSA"-
The name of the dataset, which is used for file naming convention among other things in GUIDANCE2.
seqFilter: str | None = None-
The sequence filter parameter is None, “inclusive”, or “exclusive”. If inclusive the SeqCutoff decreases for every iteration. If exclusive the SeqCutoff increases for every iteration, and so the algorithm excludes more genes from the alignment. (An OrthoEvol strategy)
columnFilter: float | None = None-
The column filter removes columns from the alignment using GUIDANCE2.
maskFilter: float | None = None-
The mask filter uses GUIDANCE2 maskLowScoresResidue script to mask the low scoring residues.
kwargs: Any = {}- The kwargs are used to configure GUIDANCE2 with specific parameters including seqCutoff and colCutoff. It can also be used to set the number of iterations and the increment number, which controls how seqCutoff and colCutoff change for each iteration.
Returns
None- Returns Guidance2 files.
pal2nal()
This Pal2Nal method works with the Pal2Nal command line wrapper.
Usage
pal2nal(
aa_alignment,
na_fasta,
output_type="paml",
nogap=True,
nomismatch=True,
downstream="paml"
)It uses a protein alignment to generate a codon alignment from the corresponding nucleic acid sequences. This is useful for downstream PAML analysis. This function also catches and removes taxa that are inconsistent with Pal2Nal’s algorithm.
Parameters
aa_alignment-
An amino acid alignment that is used as a guide for a nucleic acid alignment.
na_fasta-
The FASTA file that contains matching/ordered sequences corresponding to the aa_alignment.
output_type="paml"-
The format of the resulting alignment. (“clustal”, “paml”, “fasta”, “codon”)
nogap=True-
Removes the gaps and in-frame stop codons from the alignment to work better with PAML.
nomismatch=True-
Removes mismatched codons between protein and DNA sequences.
downstream="paml"- Used as a naming convention for a better and more obvious pipeline.
Returns
- A codon alignment.