Orthologs.Blast.BaseComparativeGenetics

Base class in the Blast module.

Usage

Source

Orthologs.Blast.BaseComparativeGenetics(
    project=None,
    project_path=os.getcwd(),
    acc_file=None,
    taxon_file=None,
    ref_species=None,
    pre_blast=False,
    post_blast=True,
    hgnc=False,
    proj_mana=None,
    copy_from_package=False,
    **kwargs
)

Methods

Name Description
__init__() This is the base class for the Blast module.
get_acc_dict() Input a list of accession numbers and return a dictionary with corresponding genes/organisms.
get_accession() Access a single accession number.
get_file_list() Turn csv column to list.
get_hgnc_gene_info() Return HGNC records in the same order as the requested symbols.
get_master_lists() Populate the organism and gene lists with a data frame.
get_orthologous_accessions() Take a single gene & return a list of accession numbers for the different orthologs.
get_orthologous_gene_sets() Access a list of accession numbers.
get_taxon_dict() Get the taxonomy information about each organism using ETE3.
get_tier_frame() Organize a dictionary by tier.

__init__()

This is the base class for the Blast module.

Usage

Source

__init__(
    project=None,
    project_path=os.getcwd(),
    acc_file=None,
    taxon_file=None,
    ref_species=None,
    pre_blast=False,
    post_blast=True,
    hgnc=False,
    proj_mana=None,
    copy_from_package=False,
    **kwargs
)

It parses an accession file in order to provide easy handling for data.

The .csv accession file contains the following header info: * “Tier” - User defined. * “Gene” - HUGO Gene Nomenclature Committee(HGNC) symbol for the genes of interest. * Query Organism - A well annotated query organism. * Other organisms - The other headers are Genus_species of other taxa.

The organisms are taken from: ftp://ftp.ncbi.nlm.nih.gov/genomes/refseq/multiprocessing/ The genes are taken from: http://www.guidetopharmacology.org/targets.jsp. The API gives the user access to their data in a higher level for downstream processing or for basic observation of the data.

Parameters
project: str | None = None

The name of the project.

project_path: str | Path | None = os.getcwd()

The location of the project, which is generally defined by the ProjectManagement configuration.

acc_file: str | None = None

The name of the accession file.

taxon_file: str | Path | None = None

A file that contains an ordered list of taxonomy ids.

ref_species: str | None = None

A reference species or organism for the blast query.

pre_blast: bool = False

A flag that gives the user access to an API that contains extra information about their genes using the mygene package.

post_blast: bool = True

A flag that is used to handle a BLAST result file, which returns information about misssing data, duplicates, etc.

hgnc: bool | str | Path = False

Enable HGNC annotation with the current complete dataset, or provide a local HGNC TSV path or alternate URL.

proj_mana: ProjectManagement | dict[str, Any] | None = None

This parameter is used to compose (vs inherit) the ProjectManagement class with the ComparativeGenetics class. This parameter allows the various blast classes to function with or without the Manager module.

copy_from_package: bool = False

Copy a packaged accession file into the project.

kwargs: Any = {}
The kwargs here are generally used for standalone blasting or for development.
Returns
None
A pandas data-frame, pivot-table, and associated lists and dictionaries.

get_acc_dict()

Input a list of accession numbers and return a dictionary with corresponding genes/organisms.

Usage

Source

get_acc_dict()
Returns
An accession dictionary who’s values are nest gene/organism lists.

get_accession()

Access a single accession number.

Usage

Source

get_accession(gene, organism)
Parameters
gene

An input gene.

organism
An input organism.
Returns
A single accession number of the target gene/organism.

get_file_list()

Turn csv column to list.

Usage

Source

get_file_list(file)
Parameters
file
Name of csv file.

get_hgnc_gene_info()

Return HGNC records in the same order as the requested symbols.

Usage

Source

get_hgnc_gene_info(gene_symbols, source=HGNC_COMPLETE_SET_URL)

get_master_lists()

Populate the organism and gene lists with a data frame.

Usage

Source

get_master_lists(df, csv_file=None)

It will also populate pre-blast attributes (mygene) and post-blast attributes (missing and duplicates) under the proper conditions.

Parameters
df

The preferred way of utilizing the function is with a data-frame.

csv_file=None
If a csv_file is given, then a data-frame will be created by reinitializing the object. (Default value = None)
Returns
An API can be utilized to access a gene list, organism list, taxon-id list, tier list/dict/data-frame, accession list/data-frame, blast query list, mygene information, and missing/duplicate information.

get_orthologous_accessions()

Take a single gene & return a list of accession numbers for the different orthologs.

Usage

Source

get_orthologous_accessions(gene)
Parameters
gene
An input gene from the accession file.
Returns
A list of accession numbers that correspond to the orthologs of the target gene.

get_orthologous_gene_sets()

Access a list of accession numbers.

Usage

Source

get_orthologous_gene_sets(go_list=None)
Parameters
go_list=None
A nested list of gene/organism lists (go_list = [[gene.1, org.1], … , [gene.n, org.n]]). (Default value = None)
Returns
An ordered list of accession numbers (or “missing”) that correspond to the go_list index.

get_taxon_dict()

Get the taxonomy information about each organism using ETE3.

Usage

Source

get_taxon_dict()
Returns
Returns several dictionaries. One is a basic organism (key) to taxonomy id (value) dictionary, and the other is a lineage dictionary with the an organism key and a lineage dictionary as the value. The lineage dictionary keys for each organism are [“class”, “family”, “genus”, “kingdowm”, “order”, “phylum”, “species”, “superkingdom”].

get_tier_frame()

Organize a dictionary by tier.

Usage

Source

get_tier_frame(tiers=None)

Each tier (key) has a value, which is a data-frame of genes associated with that tier.

Parameters
tiers=None
A list of tiers in the accession file. (Default value = None)
Returns
A nested dictionary for accessing information by tier.