---------------------------------------------------------------------- This is the API documentation for the OrthoEvol library. ---------------------------------------------------------------------- ## Ortholog discovery Infer orthologs and summarize comparative-genetics results. ## Sequence retrieval Retrieve and organize GenBank records and NCBI datasets. ## Alignment and phylogenetics Prepare alignments and run supported phylogenetic programs. ## Project and data management Create project layouts and coordinate data and database workflows. ## Execution and supporting tools Run local, PBS, Slurm, logging, and transfer helpers. ## Templates and utilities Create project templates and use shared package utilities. OrthoEvolWarning This is the OrthoEvol main warning class. The OrthoEvol-Scripts package/module will use this warning to alert uses about potential tricky code or code under development. >>> import warnings >>> from OrthoEvol import OrthoEvolWarning >>> warnings.simplefilter('ignore', OrthoEvolWarning) OrthoEvolDevelopmentWarning This is the OrthoEvol developmental code warning subclass. This warning is for alpha or beta level code which is released as part of the standard releases to mark sub-modules or functions for early adopters to test & give feedback. Code issuing this warning is likely to change or removed in a subsequent release of this package. Such code should NOT be used for production/stable code. OrthoEvolDeprecationWarning This is the Deprecation Warning subclass. This warning is for code that is no longer maintained and will be removed from this project at a later date. It may may not be working as intended, and it will not be fixed or edited. ---------------------------------------------------------------------- This is the User Guide documentation for the package. ---------------------------------------------------------------------- ### Getting Started A comparative-genetics analysis moves from candidate orthologs to sequences, alignments, and trees. OrthoEvolution coordinates these steps. External scientific programs run tasks such as BLAST searches and sequence alignment. ## Requirements OrthoEvolution supports Python 3.11 through 3.14. Before you begin, identify the external programs that your workflow uses. The Python package does not install these programs: - NCBI BLAST+ for local similarity searches - Clustal Omega, GUIDANCE2, or PAL2NAL for supported alignment workflows - PAML, PhyML, IQ-TREE, Phylip, or ETE for supported phylogenetic workflows - PBS or Slurm commands when submitting work to a cluster If your workflow uses an external program, install it before you start the analysis. The import test checks Python only. It does not test external programs or reference databases. Before you create the environment, install [`uv`](https://docs.astral.sh/uv/getting-started/installation/). ## Install from PyPI ```bash uv venv --python 3.14 .venv uv pip install --python .venv/bin/python OrthoEvol ``` ## Install from source ```bash git clone https://github.com/datasnakes/OrthoEvolution.git cd OrthoEvolution uv venv --python 3.14 .venv uv pip install --python .venv/bin/python -e . ``` ## Verify the installation ```bash .venv/bin/python -c "import OrthoEvol; print(OrthoEvol.__name__)" ``` This command makes sure that Python can import the package. It does not test external programs, network access, database contents, or a scheduler configuration. ## Choose a workflow - Start with [Ortholog Inference](02-ortholog-inference.qmd) if you have an accession table and a local BLAST database. - Use [NCBI and Sequence Retrieval](03-ncbi-sequence-retrieval.qmd) if you need an NCBI archive, BLAST database, or GenBank record. - Continue to [Alignment and Phylogenetics](04-alignment-phylogenetics.qmd) after you collect the required nucleotide or protein sequences. - Read [Project and Data Management](05-project-data-management.qmd) before you create a managed repository layout. - Browse [Work with the Example Data](08-example-data.qmd) to learn which small inputs and recorded outputs come with the repository. ### Ortholog Inference An ortholog search asks which genes in different species share an evolutionary origin. OrthoEvolution uses NCBI BLAST results to find candidates across the organisms in an accession table. It ranks sequence-similarity hits and records the selected accessions for the next analysis step. ## Before you run the search Before you run a local BLAST search, make sure that these inputs are ready: - An accession table that defines the genes and organisms - A reference species, which is normally human in the preconfigured workflow - A compatible local BLAST database - The `BLASTDB` environment variable if BLAST cannot find the database - A project directory that the workflow can write to Read the accession table before you start. Its gene labels, organism labels, and taxonomy identifiers determine how the workflow organizes each comparison. ## Prepare the accession table The accession file is a CSV table with one gene in each row. It requires the `Tier`, `Gene`, and organism columns. Use `Tier` for a group or priority that you define. | Tier | Gene | Homo_sapiens | Macaca_mulatta | Mus_musculus | Rattus_norvegicus | |---|---|---|---|---|---| | 1 | ADRA1A | NM_000680.3 | | | | | 2 | ADRA1B | NM_000679.3 | | | | Put the query organism in the first organism column. Add one query accession for each gene. Organism names use underscores, such as `Homo_sapiens`. Empty cells in the other organism columns identify accessions that the workflow will search for. Choose a query species with good annotation. Before you start, make sure that each query accession is correct. An incorrect query affects every comparison for that gene. ## Choose a BLAST method | Method | Search | Database and filter | Appropriate use | |---|---|---|---| | `1` | Local | `refseq_rna_v5` with taxonomy IDs | Multi-gene, multi-species ortholog searches | | `2` | Remote | NCBI `refseq_rna` with an Entrez query | Searches that must run through NCBI | | `None` | Local | `refseq_rna` without taxonomy IDs | A simple query, not the normal ortholog workflow | Method 1 is the recommended configuration for an ortholog search. It requires a local version 5 BLAST database. Method 2 uses an NCBI service and cannot filter with taxonomy IDs. The method controls how BLAST runs. It does not change the biological evidence required to identify an ortholog. ## Preconfigured workflow ```python from OrthoEvol.Orthologs.Blast import OrthoBlastN ortholog_search = OrthoBlastN( project="orthology-gpcr", method=1, save_data=True, acc_file="gpcr.csv", copy_from_package=True, ) ortholog_search.run() ``` This code sets up and starts the workflow. It does not install BLAST or download a database. Before you run a large analysis, make sure that BLAST can find the selected database. If `acc_file` points to your own accession table, set `copy_from_package=False`. Use the packaged-copy option only for tables that come with OrthoEvolution. ## Review the search results After the search finishes, review these output files: - Per-gene BLAST XML results - A master accession file - A table of BLAST execution times - Post-BLAST summaries when `save_data=True` The selected accessions are candidate orthologs that the computation ranks. Sequence similarity alone does not show conserved function, regulatory equivalence, or an experimentally tested evolutionary relationship. Before you interpret the biology, review all ambiguous, duplicate, and missing hits. ## When BLAST does not finish If `BLASTDB` is absent, the workflow stops early. It can stop later if BLAST is missing, a database is unavailable, or the accession table is malformed. An accession that BLAST cannot extract can also stop part of the search. Keep the log and master accession files so that you can trace a partial run. ### NCBI and Sequence Retrieval Sequence analysis starts with the correct source data. OrthoEvolution provides FTP and GenBank helpers that retrieve sequence data and prepare analysis inputs. Some downloads are large, and all remote downloads depend on NCBI services. ## Download a BLAST database ```python from pathlib import Path from OrthoEvol.Tools.ftp import NcbiFTPClient download_path = Path("databases") / "NCBI" / "blast" / "db" ncbi_ftp = NcbiFTPClient( email="researcher@example.org", max_workers=4, ) try: ncbi_ftp.getblastdb( database_name="refseq_rna", download_path=download_path, v5=True, extract=True, ) finally: ncbi_ftp.close_connection() ``` Before you start the transfer, choose a destination with enough storage. If a transfer stops, make sure that the files are complete before you use them. Version 5 selections use NCBI metadata to match the exact database. The client downloads every volume in that selection. It compares each archive with its MD5 checksum before extraction. It then keeps the checksum sidecar as an installation marker. A repeated request can reuse an installation that passed these tests. ## Download a RefSeq release selection ```python from pathlib import Path from OrthoEvol.Tools.ftp import NcbiFTPClient download_path = Path("databases") / "NCBI" / "refseq" / "release" ncbi_ftp = NcbiFTPClient( email="researcher@example.org", max_workers=4, ) try: ncbi_ftp.getrefseqrelease( collection_subset="vertebrate_mammalian", seqtype="rna", seqformat="gbff", download_path=download_path, extract=True, ) finally: ncbi_ftp.close_connection() ``` This request downloads RNA GenBank flat files for the vertebrate-mammalian collection. A local marker records the NCBI release number and selected files. The client skips a later request only if the marker and output files match the current release. Common sequence formats include: | Extension | Typical content | |---|---| | `.fna` | Genomic or other nucleotide FASTA sequences | | `.ffn` | Nucleotide FASTA sequences for coding regions | | `.faa` | Protein FASTA sequences | | `.frn` | Non-coding RNA FASTA sequences | | `.gbff` | GenBank flat-file records with sequence annotations | Select `seqtype` and `seqformat` for the exact NCBI collection that you need. A valid file format does not show that the release, molecule type, or taxa fit your analysis. ## GenBank processing The `GenBank` class connects a project to local GenBank archives and BioSQL resources. GenBank records contain sequences and their annotations. The class can: - Locate accessions in a local GenBank database - Write individual GenBank records - Extract nucleotide or protein FASTA files - Organize outputs by gene and organism Before you create the class, make sure that the project paths and required database resources exist. Several methods use NCBI database locations from a managed repository. ## Record what you downloaded Record the NCBI database name and its release or retrieval date. Also record the available file checksums and the accession-table version. NCBI content can change without an OrthoEvolution release. The package version alone cannot reproduce a retrieval. ## When retrieval does not complete First identify which part of the retrieval stopped. The cause can be a network interruption, a missing remote file, insufficient storage, a checksum mismatch, or an incomplete BioSQL setup. Do not use a partial archive. A file at the destination is not proof of a complete transfer. ### Alignment and Phylogenetics An alignment places comparable sequence positions in the same columns. A phylogenetic analysis then uses that alignment to estimate evolutionary relationships. OrthoEvolution coordinates both stages, but external scientific programs perform the computations. ## Alignment workflow `MultipleSequenceAlignment` can coordinate Clustal Omega, GUIDANCE2, and PAL2NAL. Select the sequence type before you align the records: - Use nucleotide inputs where a nucleotide alignment is required - Use amino-acid inputs for protein alignment and filtering - Retain the coding-sequence relationship when PAL2NAL projects a protein alignment back to codons GUIDANCE2 writes separate files for retained sequences, removed sequences, alignments, filtered columns, and masked alignments. Record which filtering path produced the input for the next step. ## Run a Clustal Omega alignment ```python from pathlib import Path from OrthoEvol.Orthologs.Align import ClustalO input_fasta = Path("examples/example-data/Alignment_Filter/HTR1A.faa") output_fasta = Path("HTR1A_aligned.fasta") alignment = ClustalO( infile=input_fasta, outfile=output_fasta, outfmt="fasta", ) alignment.runclustalomega() ``` `ClustalO` builds and runs the external `clustalo` command. Before you run it, make sure that `clustalo` is on `PATH`. The input file must contain one biological sequence type. Use proteins if the alignment will guide a later PAL2NAL codon alignment. ## Configure GUIDANCE2 ```python from OrthoEvol.Orthologs.Align import Guidance2Commandline guidance = Guidance2Commandline( seqFile="examples/example-data/Alignment_Filter/HTR1A.faa", msaProgram="MAFFT", seqType="aa", outDir="HTR1A_guidance2", ) ``` This object records the input, alignment program, sequence type, and output directory. To run the command, you still need GUIDANCE2 and the selected alignment program. Review the retained and removed sequences. Filtering does not always improve the biological analysis. ## Convert an alignment for PhyML ```python from OrthoEvol.Orthologs.Phylogenetics import PhyML, RelaxPhylip RelaxPhylip("HTR1A_aligned.fasta", "HTR1A_aligned.phy") phyml = PhyML("HTR1A_aligned.phy", datatype="aa") phyml.run(model="WAG", alpha="e", bootstrap=100) ``` `RelaxPhylip` converts the alignment to the relaxed PHYLIP format that the wrapper reads. PhyML is optional external software. OrthoEvolution does not install it. If the `phyml` executable is absent, `PhyML` raises `FileNotFoundError` when you create the object. ## Phylogenetic tools The package provides interfaces for PAML, PhyML, IQ-TREE, Phylip, ETE, and tree visualization. Each interface has requirements for its executable, input format, and operating system. Before you automate one, read its constructor and method details in the API reference. The bundled Phylip wrapper runs only on Linux. The repository keeps the ETE3/PAML and filtered-tree interfaces for existing workflows. It does not include a complete portable tutorial for their external programs. Their reference pages describe the API, but they do not test your local installation. ## Interpret the tree The sampled taxa, selected sequences, alignment, filters, substitution model, and program configuration all affect an inferred tree. A clear plot does not show that these choices fit the biological question. Keep the unaligned sequences, final alignment, program configuration, logs, and tree files together. ## Before trusting a result If a command stops immediately, make sure that its executable and input format are correct. A completed command can still give a misleading result. This can happen when identifiers differ, aligned regions are not homologous, or filters remove the variation required for inference. ### Project and Data Management A comparative-genetics project can produce many files across several analysis steps. The `Manager` package gives repository, user, project, database, raw-data, and web directories a consistent hierarchy. Its classes can create these directories from bundled Cookiecutter templates. ## Create a project layout ```python from OrthoEvol.Manager.management import ProjectManagement project_manager = ProjectManagement( repo="comparative-genetics", user="researcher", project="receptor-family", research=None, research_type="comparative_genetics", new_project=True, ) ``` This code writes directories and files from a template. Before you set `new_project=True`, use an explicit destination and make sure that it is the intended location. ## See the managed directory layout The management classes and templates use the repository hierarchy below. A constructor creates only the part selected by `new_repo`, `new_user`, `new_project`, `new_db`, `new_research`, or `new_app`. Some paths exist in the manager before a tool writes data to their directories. ```text / ├── docs/ ├── users/ │ └── / │ ├── archive/ │ ├── databases/ │ │ ├── / │ │ ├── ITIS/ │ │ └── NCBI/ │ │ ├── blast/ │ │ │ ├── db/ │ │ │ ├── seqidlists/ │ │ │ └── windowmasker_files/ │ │ ├── pub/ │ │ │ └── taxonomy/ │ │ └── refseq/ │ │ └── release/ │ ├── index/ │ ├── log/ │ ├── manuscripts/ │ ├── other/ │ └── projects/ │ └── / │ ├── archive/ │ └── / │ └── / │ ├── data/ │ ├── index/ │ ├── raw_data/ │ └── web/ │ └── / └── web/ ├── flask/ ├── ftp/ ├── shiny/ └── wasabi/ ``` The user's `databases/` directory contains project databases. They do not live inside `projects//`. The manager adds a research directory only when you supply both a research type and a research name. ## Choose what to create | Option | Created area | |---|---| | `new_repo=True` | Repository-level directories and templates | | `new_user=True` | A user's archive, database, project, and support directories | | `new_db=True` | Database directories selected by the database configuration | | `new_project=True` | The named project beneath the user's project hierarchy | | `new_research=True` | The named research directory beneath its research type | | `new_app=True` | The selected web application beneath newly created research | These flags create content. They do not search an existing layout. Before you use a flag, supply the parent identifiers for that level. Make sure that the resolved paths are correct before you enable more than one creation flag. The current `ProjectManagement` implementation handles `new_app=True` inside the `new_research=True` branch. You must also supply `research`, `research_type`, and `app`. ## Follow the project hierarchy - `Management` maps package resources and an optional repository root. - `RepoManagement` adds repository-wide user, documentation, and web paths. - `UserManagement` adds user-specific databases, archives, projects, and raw data. - `ProjectManagement` specializes the hierarchy for a research project. - `DataMana` and the database-management classes dispatch configured data and database operations. ## Review the configuration A managed workflow can combine Python objects with YAML configuration. Before you dispatch the workflow, make sure that the YAML structure and referenced paths are correct. Valid YAML can still name an invalid database strategy, an unavailable path, or an unsupported operation. ## Practice in a disposable project Start with a small disposable project while you learn the hierarchy. Make sure that each manager supplies the paths required by the next class. This practice reduces the risk of writing to the wrong directory or starting a large data operation with the wrong configuration. ### Cluster Execution A cluster scheduler assigns computing resources to jobs. OrthoEvolution includes PBS helpers and a small synchronous Slurm client. These interfaces call scheduler commands on the execution host. They do not install a scheduler or translate the resource rules for your site. ## Slurm example ```python from pathlib import Path from OrthoEvol.Tools.slurm import SlurmClient slurm = SlurmClient() job_id = slurm.submit(Path("analysis.sbatch")) active_jobs = slurm.active_jobs() job_history = slurm.job_history(job_id) ``` `submit()` sends an existing batch script and returns its job identifier. `active_jobs()` requests the active jobs for one user. `job_history()` requests the allocation record for one job. The client returns `SlurmJob` objects and keeps the state and node information from Slurm. ## PBS support The PBS package provides `Qsub` for job submission and `Qstat` for job status. PBS output formats differ between systems. Before you automate these interfaces, make sure that your local output matches the parser format. ```python from OrthoEvol.Tools.pbs import BaseQstat, BaseQsub pbs_job = BaseQsub( job_name="orthologs", pbs_script="analysis.pbs", pbs_working_dir="pbs-jobs", ) pbs_job.copy_supplied_script( supplied_script=pbs_job.supplied_pbs_script, new_script=pbs_job.pbs_script, ) pbs_job.submit_pbs_script() job_status = BaseQstat( job_id=pbs_job.pbs_job_id, home="pbs-status", ) job_status.run_qstat(csv_flag=True) ``` `BaseQsub` creates a unique work directory. It copies the supplied script into that directory before submission. `BaseQstat` records one full status response. This example runs only on a host with compatible `qsub` and `qstat` commands. Before you create the status client, make sure that `pbs_job.pbs_job_id` has a value. A failed submission leaves it unset. ## Adapt the job to your cluster Your cluster defines its partitions, queues, accounts, time limits, memory syntax, modules, and storage paths. Put these values in the batch script or site configuration. Do not assume that an example from another cluster will run without changes. ## Diagnose scheduler errors If a scheduler command is missing, the request stops immediately. Other failures start after submission or status inspection. Common causes include an invalid account, unavailable resources, queue rules, missing accounting data, and malformed scheduler output. Save the exact error and job identifier for the cluster administrator. ### Troubleshooting OrthoEvolution connects Python, external programs, data files, network services, and cluster schedulers. First identify which system stopped. A Python traceback can report a failure that started outside Python. ## Confirm the environment ```bash .venv/bin/python --version .venv/bin/python -c "import OrthoEvol; print(OrthoEvol.__name__)" .venv/bin/python -m pip check ``` Use Python 3.11 through 3.14. Run each command with the same local environment that contains OrthoEvolution. ## Classify the failure | Boundary | First checks | |---|---| | Python import | Interpreter path, installed package, dependency conflicts | | BLAST | Executable on `PATH`, `BLASTDB`, database files, accession input | | Alignment | Executable path, sequence type, file format, writable output | | Phylogenetics | Program installation, model settings, alignment validity | | NCBI retrieval | Network access, remote path, storage, archive validation | | PBS or Slurm | Scheduler command, account, queue or partition, job identifier | ## Preserve useful evidence Save the exact command, traceback, log, configuration file, and smallest input that fails. Also record the package, Python, and external-program versions. Before you share the report, remove credentials and private paths. ### Work with the Example Data The repository includes small files for ortholog searches, sequences, alignments, and trees. Use them to learn each file format or practice one operation at a time. You do not need to start with a large dataset. These files are historical test fixtures and recorded outputs. They do not form a complete analysis or a current biological reference dataset. ## Find the files ```text examples/example-data/ ├── BLASTest_MAF.csv ├── BLASTest_TIME.csv ├── BLASTest_mygene.csv ├── HTR1A_aligned.phy ├── HTR1A_aligned_cds_nucl.fasta ├── MASTER_HTR1A_CDS1.ffn ├── example_pba.xlsx ├── organisms.csv ├── species_tree.nw └── Alignment_Filter/ ├── HTR1A.faa ├── HTR1A.ffn ├── HTR1A_G2.faa ├── HTR1A_G2.ffn ├── HTR1A_G2_aa.aln ├── HTR1A_G2_removed.ffn ├── HTR1A_P2N_na.aln ├── AA_Guidance2/ └── NA_Guidance2/ ``` The GUIDANCE2 directories contain intermediate scores, alignments, and filter files. Keep them with the input and final alignment. Together, these files show which sequences or columns the filter removed. ## Choose an example by task | Task | Start with | What it demonstrates | |---|---|---| | Inspect accession tables | `BLASTest_mygene.csv` | Gene-by-organism accession layout | | Review BLAST summaries | `BLASTest_MAF.csv`, `BLASTest_TIME.csv` | Recorded accession and timing outputs | | Practice protein alignment | `Alignment_Filter/HTR1A.faa` | Multi-sequence protein FASTA input | | Inspect coding sequences | `Alignment_Filter/HTR1A.ffn` | Matching coding-region nucleotide FASTA | | Inspect a filtered alignment | `Alignment_Filter/HTR1A_G2_aa.aln` | Recorded GUIDANCE2-derived alignment | | Inspect PAL2NAL output | `Alignment_Filter/HTR1A_P2N_na.aln` | Protein-guided nucleotide alignment | | Read PHYLIP input | `HTR1A_aligned.phy` | Alignment in PHYLIP format | | Read a tree | `species_tree.nw` | Newick-formatted tree | | Inspect the organism set | `organisms.csv` | One organism name per row | You can find these files in the repository's [`examples/example-data`](https://github.com/datasnakes/OrthoEvolution/tree/main/examples/example-data) directory. ## Inspect a sequence file ```python from pathlib import Path from Bio import SeqIO sequence_file = Path("examples/example-data/Alignment_Filter/HTR1A.faa") records = list(SeqIO.parse(sequence_file, "fasta")) assert records assert all(record.seq for record in records) ``` This test makes sure that Biopython can read at least one non-empty FASTA record. It does not test orthology, sequence quality, taxonomic coverage, or the relationship between the protein and nucleotide files. ## Use the examples as snapshots Before a tool changes an example, copy the file to a disposable work directory. Record the example name and the program version for each new result. The bundled accession values, alignment filters, and trees are snapshots. Do not treat them as current data.