utilities.BlastUtils
Usage
utilities.BlastUtils()Methods
| Name | Description |
|---|---|
| __init__() | Various utilities to help with blast specific functionality. |
| accession_csv2sqlite() | Convert am OrthoEvol csv accession file to an sqlite3 database. |
| accession_sqlite2pandas() | Convert a sqlite3 database with an OrthoEvol accession table to a pandas dataframe. |
| analyze_duplicate_accessions() | Classify duplicate accessions and calculate report counts once. |
| gene_list_config() | Create or use a blast configuration file (accession file). |
| get_dup_acc() | Get duplicated accession numbers during post-blast analysis. |
| get_miss_acc() | Get missing accession numbers during post-blast analysis. |
| map_func() | Format/parse hit ids generated from blast xml results. |
| my_gene_info() | Use Biothings’ MyGene api to get information about genes. |
| paml_org_formatter() | Take a list of organisms and format each organism name for PAML, |
__init__()
Various utilities to help with blast specific functionality.
Usage
__init__()accession_csv2sqlite()
Convert am OrthoEvol csv accession file to an sqlite3 database.
Usage
accession_csv2sqlite(acc_file, table_name, db_name, path)Parameters
acc_file: str-
The name of the accession file. The file name is used to create a table in the sqlite3 database. Any periods will be replaced with underscores.
table_name: str-
The name of the table in the database.
db_name: str-
The name of the new database.
path: str- The relative path of the csv file and the database.
accession_sqlite2pandas()
Convert a sqlite3 database with an OrthoEvol accession table to a pandas dataframe.
Usage
accession_sqlite2pandas(table_name, db_name, path, exists=True, acc_file=None)Parameters
table_name: str-
Name of the table in the database.
db_name: str-
The name of the new database.
path: str-
The relative path of the csv file and the database.
exists: bool = True-
A flag used to create a database if needed.
acc_file: str | None = None- The name of the accession file. The file name is used to create a table in the sqlite3 database. Any periods will be replaced with underscores.
Returns
pd.DataFrame- A pandas DataFrame containing the accession data.
analyze_duplicate_accessions()
Classify duplicate accessions and calculate report counts once.
Usage
analyze_duplicate_accessions(acc_dict, gene_list, org_list)gene_list_config()
Create or use a blast configuration file (accession file).
Usage
gene_list_config(file, data_path, gene_list, taxon_dict, logger)This function configures different files for new BLASTS. It also helps recognize whether or not a BLAST was terminated in the middle of the workflow. This removes the last line of the accession file if it is incomplete.
Parameters
file: str.-
An accession file to analyze.
data_path: str.-
The path of the accession file.
gene_list: list.-
A gene list in the same order as the accession file.
taxon_dict: dict.-
A taxon id dictionary for logging purposes.
logger: LogIt.- A LogIt logger for logging.
Returns
- Returns a continued gene_list to pick up from an interrupted Blast.
get_dup_acc()
Get duplicated accession numbers during post-blast analysis.
Usage
get_dup_acc(acc_dict, gene_list, org_list)Parameters
acc_dict: Mapping[str, Sequence[Sequence[str]]]-
A dictionary with accession numbers as keys, and a gene/organism list as values.
gene_list: Sequence[str]-
A full list of genes.
org_list: Sequence[str]- A full list of organisms.
Returns
dict.- A master duplication dictionary used to initialize the duplicate class variables.
get_miss_acc()
Get missing accession numbers during post-blast analysis.
Usage
get_miss_acc(acc_dataframe)Parameters
acc_dataframe: pd.DataFrame- A pandas dataframe containing the accession csv file data(post BLAST).
Returns
dict.- A dictionary with data about the missing accession numbers by Gene and by Organism.
map_func()
Format/parse hit ids generated from blast xml results.
Usage
map_func(hit)Parameters
hit: Bio.SearchIO.HSP | Bio.SearchIO.Hit- A BLAST hit object from SearchIO.
Returns
Bio.SearchIO.HSP | Bio.SearchIO.Hit- The hit object with formatted id, id1 (accession), and id2 (gi).
my_gene_info()
Use Biothings’ MyGene api to get information about genes.
Usage
my_gene_info(acc_dataframe, blast_query="Homo_sapiens")Parameters
acc_dataframe: pd.DataFrame.-
A pandas dataframe containing the accession csv file data.
blast_query: str. = "Homo_sapiens"- The query organism for used during Blasting.
Returns
pd.DataFrame.- Returns a data-frame with hot data about each gene.
paml_org_formatter()
Take a list of organisms and format each organism name for PAML,
Usage
paml_org_formatter(organisms)which can only take names that are less than a certain length (36 characters?).
Parameters
organisms: list- A list of organisms.
Returns
list- A list of formatted organism names for PAML.