Orthologs.GenBank.GenBank

This class will handle GenBank files in various ways.

Usage

Source

Orthologs.GenBank.GenBank(
    project,
    project_path=None,
    solo=False,
    multi=True,
    archive=False,
    min_fasta=True,
    blast=OrthoBlastN,
    **kwargs
)

Methods

Name Description
__init__() Handle GenBank files in various ways.
create_post_blast_gbk_records() Create a single GenBank file for each ortholog.
gbk_quality_control() Ensures the quality or validity of the retrieved genbank record.
gbk_upload() Upload a BioSQL database with target GenBank data (.gbk files).
get_fasta_files() Create FASTA files for each GenBank record in the accession dictionary.
get_gbk_file() Search a GenBank database for a target accession number.
multi_fasta() Append an othologous sequence of a feature to a uniquely named file.
name_fasta_file() Provide a uniquely named FASTA file.
protein_gi_fetch() Retrieve the protein gi number.
solo_fasta() This method writes a sequence of a feature to a uniquely named file using a dictionary for formatting.
write_fasta_files() Create a dictionary for formatting the FASTA header & sequence.

__init__()

Handle GenBank files in various ways.

Usage

Source

__init__(
    project,
    project_path=None,
    solo=False,
    multi=True,
    archive=False,
    min_fasta=True,
    blast=OrthoBlastN,
    **kwargs
)

It allows for refseq-release .gbff files to be downloaded from NCBI and uploaded to a BioSQL database (biopython). Single .gbk files can be downloaded from the .gbff, and uploaded to a custom BopSQL database for faster acquisition of GenBank data.

Parameters
project: str

The name of the project.

project_path: str | Path | None = None

The relative path to the project.

solo: bool = False

A flag for adding single fasta files.

multi: bool = True

A flag for adding multi-fasta files.

archive: bool = False

A flag for archiving current GenBank Data.

min_fasta: bool = True

A flag for minimizing FASTA file headers.

blast: class | dict | None = OrthoBlastN

The blast parameter is used for composing various Orthologs.Blast classes. Can be a class, a dict, or None.

kwargs: dict = {}
Additional keyword arguments for configuration.
Returns
.gbff files/databases, .gbk files/databases, & FASTA files.

create_post_blast_gbk_records()

Create a single GenBank file for each ortholog.

Usage

Source

create_post_blast_gbk_records(org_list, gene_dict)

After a blast has completed and the accession numbers have been compiled into an accession file, this class searches a local NCBI refseq release database composed of GenBank records. This method will create a single GenBank file (.gbk) for each ortholog with an accession number. The create_post_blast_gbk_records is only callable if the the instance is composed by one of the Blast classes. This method also requires an NCBI refseq release database to be set up with the proper GenBank Flat Files (.gbff) files.

Parameters
org_list

List of organisms

gene_dict
A nested dictionary for accessing accession numbers. (e.g. gene_dict[GENE][ORGANISM} yields an accession number)
Returns
Does not return an object, but creates genbank files.

gbk_quality_control()

Ensures the quality or validity of the retrieved genbank record.

Usage

Source

gbk_quality_control(gbk_file, gene, organism)

It takes the GenBank record and check to make sure the Gene and Organism from the GenBank record match the Gene and Organism from the accession file. If not, then the Blast has returned the wrong accession number.

:param gbk_file:  The path to a GenBank file.
:param gene:  A gene name from the Accession file.
:param organism:  A gene name from the Accession file.
:return:

gbk_upload()

Upload a BioSQL database with target GenBank data (.gbk files).

Usage

Source

gbk_upload()

This method is only usable after creating GenBank records with this class. It uploads a BioSQL databases with target GenBank data (.gbk files). This creates a compact set of data for each project.

Returns
Does not return an object.

get_fasta_files()

Create FASTA files for each GenBank record in the accession dictionary.

Usage

Source

get_fasta_files(acc_dict, db=True)

It can search through a BioSQL database or it can crawl a directory for .gbk files.

Parameters
acc_dict

An accession dictionary like the one created by CompGenObjects.

db=True
A flag that determines whether or not to use the custom BioSQL database or to use .gbk files. (Default value = True)
Returns
Returns FASTA files for each GenBank record.

get_gbk_file()

Search a GenBank database for a target accession number.

Usage

Source

get_gbk_file(accession, gene, organism, server_flag=None)

This function searches through the given NCBI databases (created by uploading NCBI refseq .gbff files to a BioPython BioSQL database) and creates single GenBank files. This function can be used after a blast or on its own. If used on it’s own then the NCBI .db files must be manually moved to the proper directories.

Parameters
accession

Accession number of interest without the version.

gene

Target gene of the accession number parameter.

organism

Target organism of the accession number parameter.

server_flag=None
(Default value = None)
Returns

multi_fasta()

Append an othologous sequence of a feature to a uniquely named file.

Usage

Source

multi_fasta(na_entry, aa_entry, fmt)

Usese a dictionary for formatting.

Parameters
na_entry

A string representing the Nucleic Acid sequence data in FASTA format.

aa_entry

A string representing the Amino Acid sequence data in FASTA format.

fmt
A dictionary for formatting the FASTA entries and the file names.
Returns
Does not return an object, but creates or appends to a multi entry FASTA file.

name_fasta_file()

Provide a uniquely named FASTA file.

Usage

Source

name_fasta_file(path, gene, org, feat_type, feat_type_rank, extension, mode)
  • Coding sequence:
    • Single - “/_.”
    • Multi - “/.”
  • Other:
    • Single - “/.”
    • Multi - “/_.”
Parameters
path: str | Path

The path where the file will be made.

gene: str

The gene name.

org: str

The organism name.

feat_type: str

The type of feature from the GenBank record (CDS, UTR, misc_feature, variation, etc.).

feat_type_rank: str

The feature type + the rank (There can be multiple misc_features and variations).

extension: str

The file extension (“.ffn”, “.faa”, “.fna”, “.fasta”).

mode: str
The mode (“w” or “a”) for writing the file. Write to a solo-FASTA file. Append a multi-FASTA file.
Returns
file object
The uniquely named FASTA file.

protein_gi_fetch()

Retrieve the protein gi number.

Usage

Source

protein_gi_fetch(feature)
Parameters
feature
Search the protein feature for the GI number.
Returns
The protein GI number as a string.

solo_fasta()

This method writes a sequence of a feature to a uniquely named file using a dictionary for formatting.

Usage

Source

solo_fasta(na_entry, aa_entry, fmt)
Parameters
na_entry

A string representing the Nucleic Acid sequence data in FASTA format.

aa_entry

A string representing the Amino Acid sequence data in FASTA format.

fmt
A dictionary for formatting the FASTA entries and the file names.
Returns
Does not return an object, but creates single entry FASTA files.

write_fasta_files()

Create a dictionary for formatting the FASTA header & sequence.

Usage

Source

write_fasta_files(record, acc_dict)
Parameters
record

A GenBank record created by BioPython.

acc_dict
Accession dictionary from the CompGenObjects class.
Returns