Orthologs.GenBank.GenBank
This class will handle GenBank files in various ways.
Usage
Orthologs.GenBank.GenBank(
project,
project_path=None,
solo=False,
multi=True,
archive=False,
min_fasta=True,
blast=OrthoBlastN,
**kwargs
)Methods
| Name | Description |
|---|---|
| __init__() | Handle GenBank files in various ways. |
| create_post_blast_gbk_records() | Create a single GenBank file for each ortholog. |
| gbk_quality_control() | Ensures the quality or validity of the retrieved genbank record. |
| gbk_upload() | Upload a BioSQL database with target GenBank data (.gbk files). |
| get_fasta_files() | Create FASTA files for each GenBank record in the accession dictionary. |
| get_gbk_file() | Search a GenBank database for a target accession number. |
| multi_fasta() | Append an othologous sequence of a feature to a uniquely named file. |
| name_fasta_file() | Provide a uniquely named FASTA file. |
| protein_gi_fetch() | Retrieve the protein gi number. |
| solo_fasta() | This method writes a sequence of a feature to a uniquely named file using a dictionary for formatting. |
| write_fasta_files() | Create a dictionary for formatting the FASTA header & sequence. |
__init__()
Handle GenBank files in various ways.
Usage
__init__(
project,
project_path=None,
solo=False,
multi=True,
archive=False,
min_fasta=True,
blast=OrthoBlastN,
**kwargs
)It allows for refseq-release .gbff files to be downloaded from NCBI and uploaded to a BioSQL database (biopython). Single .gbk files can be downloaded from the .gbff, and uploaded to a custom BopSQL database for faster acquisition of GenBank data.
Parameters
project: str-
The name of the project.
project_path: str | Path | None = None-
The relative path to the project.
solo: bool = False-
A flag for adding single fasta files.
multi: bool = True-
A flag for adding multi-fasta files.
archive: bool = False-
A flag for archiving current GenBank Data.
min_fasta: bool = True-
A flag for minimizing FASTA file headers.
blast: class | dict | None = OrthoBlastN-
The blast parameter is used for composing various Orthologs.Blast classes. Can be a class, a dict, or None.
kwargs: dict = {}- Additional keyword arguments for configuration.
Returns
- .gbff files/databases, .gbk files/databases, & FASTA files.
create_post_blast_gbk_records()
Create a single GenBank file for each ortholog.
Usage
create_post_blast_gbk_records(org_list, gene_dict)After a blast has completed and the accession numbers have been compiled into an accession file, this class searches a local NCBI refseq release database composed of GenBank records. This method will create a single GenBank file (.gbk) for each ortholog with an accession number. The create_post_blast_gbk_records is only callable if the the instance is composed by one of the Blast classes. This method also requires an NCBI refseq release database to be set up with the proper GenBank Flat Files (.gbff) files.
Parameters
org_list-
List of organisms
gene_dict- A nested dictionary for accessing accession numbers. (e.g. gene_dict[GENE][ORGANISM} yields an accession number)
Returns
- Does not return an object, but creates genbank files.
gbk_quality_control()
Ensures the quality or validity of the retrieved genbank record.
Usage
gbk_quality_control(gbk_file, gene, organism)It takes the GenBank record and check to make sure the Gene and Organism from the GenBank record match the Gene and Organism from the accession file. If not, then the Blast has returned the wrong accession number.
:param gbk_file: The path to a GenBank file.
:param gene: A gene name from the Accession file.
:param organism: A gene name from the Accession file.
:return:
gbk_upload()
Upload a BioSQL database with target GenBank data (.gbk files).
Usage
gbk_upload()This method is only usable after creating GenBank records with this class. It uploads a BioSQL databases with target GenBank data (.gbk files). This creates a compact set of data for each project.
Returns
- Does not return an object.
get_fasta_files()
Create FASTA files for each GenBank record in the accession dictionary.
Usage
get_fasta_files(acc_dict, db=True)It can search through a BioSQL database or it can crawl a directory for .gbk files.
Parameters
acc_dict-
An accession dictionary like the one created by CompGenObjects.
db=True- A flag that determines whether or not to use the custom BioSQL database or to use .gbk files. (Default value = True)
Returns
- Returns FASTA files for each GenBank record.
get_gbk_file()
Search a GenBank database for a target accession number.
Usage
get_gbk_file(accession, gene, organism, server_flag=None)This function searches through the given NCBI databases (created by uploading NCBI refseq .gbff files to a BioPython BioSQL database) and creates single GenBank files. This function can be used after a blast or on its own. If used on it’s own then the NCBI .db files must be manually moved to the proper directories.
Parameters
accession-
Accession number of interest without the version.
gene-
Target gene of the accession number parameter.
organism-
Target organism of the accession number parameter.
server_flag=None- (Default value = None)
Returns
multi_fasta()
Append an othologous sequence of a feature to a uniquely named file.
Usage
multi_fasta(na_entry, aa_entry, fmt)Usese a dictionary for formatting.
Parameters
na_entry-
A string representing the Nucleic Acid sequence data in FASTA format.
aa_entry-
A string representing the Amino Acid sequence data in FASTA format.
fmt- A dictionary for formatting the FASTA entries and the file names.
Returns
- Does not return an object, but creates or appends to a multi entry FASTA file.
name_fasta_file()
Provide a uniquely named FASTA file.
Usage
name_fasta_file(path, gene, org, feat_type, feat_type_rank, extension, mode)- Coding sequence:
- Single - “
/ _ . ” - Multi - “
/ . ”
- Single - “
- Other:
- Single - “
/ . ” - Multi - “
/ _ . ”
- Single - “
Parameters
path: str | Path-
The path where the file will be made.
gene: str-
The gene name.
org: str-
The organism name.
feat_type: str-
The type of feature from the GenBank record (CDS, UTR, misc_feature, variation, etc.).
feat_type_rank: str-
The feature type + the rank (There can be multiple misc_features and variations).
extension: str-
The file extension (“.ffn”, “.faa”, “.fna”, “.fasta”).
mode: str- The mode (“w” or “a”) for writing the file. Write to a solo-FASTA file. Append a multi-FASTA file.
Returns
file object- The uniquely named FASTA file.
protein_gi_fetch()
Retrieve the protein gi number.
Usage
protein_gi_fetch(feature)Parameters
feature- Search the protein feature for the GI number.
Returns
- The protein GI number as a string.
solo_fasta()
This method writes a sequence of a feature to a uniquely named file using a dictionary for formatting.
Usage
solo_fasta(na_entry, aa_entry, fmt)Parameters
na_entry-
A string representing the Nucleic Acid sequence data in FASTA format.
aa_entry-
A string representing the Amino Acid sequence data in FASTA format.
fmt- A dictionary for formatting the FASTA entries and the file names.
Returns
- Does not return an object, but creates single entry FASTA files.
write_fasta_files()
Create a dictionary for formatting the FASTA header & sequence.
Usage
write_fasta_files(record, acc_dict)Parameters
record-
A GenBank record created by BioPython.
acc_dict- Accession dictionary from the CompGenObjects class.