{"created":"2023-06-26T11:00:52.641842+00:00","id":1471,"links":{},"metadata":{"_buckets":{"deposit":"1d08c15e-1651-484c-8837-362e243fd462"},"_deposit":{"created_by":29,"id":"1471","owners":[29],"pid":{"revision_id":0,"type":"depid","value":"1471"},"status":"published"},"_oai":{"id":"oai:oist.repo.nii.ac.jp:00001471","sets":["6:45"]},"author_link":["8909","8910"],"item_10001_biblio_info_7":{"attribute_name":"Bibliographic Information","attribute_value_mlt":[{"bibliographicIssueDates":{"bibliographicIssueDate":"2020-02-06","bibliographicIssueDateType":"Issued"},"bibliographicIssueNumber":"1","bibliographicPageStart":"48","bibliographicVolumeNumber":"21","bibliographic_titles":[{},{"bibliographic_title":"BMC Bioinformatics","bibliographic_titleLang":"en"}]}]},"item_10001_creator_3":{"attribute_name":"Author","attribute_type":"creator","attribute_value_mlt":[{"creatorNames":[{"creatorName":"Gao, Kun"}],"nameIdentifiers":[{}]},{"creatorNames":[{"creatorName":"Miller, Jonathan"}],"nameIdentifiers":[{}]}]},"item_10001_description_5":{"attribute_name":"Abstract","attribute_value_mlt":[{"subitem_description":"Background\nThe evolutionary history of genes serves as a cornerstone of contemporary biology. Most conserved sequences in mammalian genomes don't code for proteins, yielding a need to infer evolutionary history of sequences irrespective of what kind of functional element they may encode. Thus, sequence-, as opposed to gene-, centric modes of inferring paths of sequence evolution are increasingly relevant. Customarily, homologous sequences derived from the same direct ancestor, whose ancestral position in two genomes is usually conserved, are termed \"primary\" (or \"positional\") orthologs. Methods based solely on similarity don't reliably distinguish primary orthologs from other homologs; for this, genomic context is often essential. Context-dependent identification of orthologs traditionally relies on genomic context over length scales characteristic of conserved gene order or whole-genome sequence alignment, and can be computationally intensive. \n\nResults\nWe demonstrate that short-range sequence context-as short as a single \"maximal\" match- distinguishes primary orthologs from other homologs across whole genomes. On mammalian whole genomes not preprocessed by repeat-masker, potential orthologs are extracted by genome intersection as \"non-nested maximal matches:\" maximal matches that are not nested into other maximal matches. It emerges that on both nucleotide and gene scales, non-nested maximal matches recapitulate primary or positional orthologs with high precision and high recall, while the corresponding computation consumes less than one thirtieth of the computation time required by commonly applied whole-genome alignment methods. In regions of genomes that would be masked by repeat-masker, non-nested maximal matches recover orthologs that are inaccessible to Lastz net alignment, for which repeat-masking is a prerequisite. mmRBHs, reciprocal best hits of genes containing non-nested maximal matches, yield novel putative orthologs, e.g. around 1000 pairs of genes for human-chimpanzee. \n\nConclusions\nWe describe an intersection-based method that requires neither repeat-masking nor alignment to infer evolutionary history of sequences based on short-range genomic sequence context. Ortholog identification based on non-nested maximal matches is parameter-free, and less computationally intensive than many alignment-based methods. It is especially suitable for genome-wide identification of orthologs, and may be applicable to unassembled genomes. We are agnostic as to the reasons for its effectiveness, which may reflect local variation of mean mutation rate.","subitem_description_type":"Other"}]},"item_10001_publisher_8":{"attribute_name":"Publisher","attribute_value_mlt":[{"subitem_publisher":"BioMed Central Ltd"}]},"item_10001_relation_13":{"attribute_name":"PubMedNo.","attribute_value_mlt":[{"subitem_relation_type":"isIdenticalTo","subitem_relation_type_id":{"subitem_relation_type_id_text":"info:pmid/32028880","subitem_relation_type_select":"PMID"}}]},"item_10001_relation_14":{"attribute_name":"DOI","attribute_value_mlt":[{"subitem_relation_type":"isIdenticalTo","subitem_relation_type_id":{"subitem_relation_type_id_text":"info:doi/10.1186/s12859-020-3384-2","subitem_relation_type_select":"DOI"}}]},"item_10001_relation_16":{"attribute_name":"情報源","attribute_value_mlt":[{"subitem_relation_name":[{"subitem_relation_name_text":"https://creativecommons.org/licenses/by/4.0/"}]}]},"item_10001_relation_17":{"attribute_name":"Related site","attribute_value_mlt":[{"subitem_relation_type_id":{"subitem_relation_type_id_text":"https://doi.org/10.1186/s12859-020-3384-2","subitem_relation_type_select":"DOI"}}]},"item_10001_rights_15":{"attribute_name":"Rights","attribute_value_mlt":[{"subitem_rights":"© 2020 The Author(s). "}]},"item_10001_source_id_9":{"attribute_name":"ISSN","attribute_value_mlt":[{"subitem_source_identifier":"1471-2105","subitem_source_identifier_type":"ISSN"}]},"item_10001_version_type_20":{"attribute_name":"Author's flag","attribute_value_mlt":[{"subitem_version_resource":"http://purl.org/coar/version/c_970fb48d4fbd8a85","subitem_version_type":"VoR"}]},"item_files":{"attribute_name":"ファイル情報","attribute_type":"file","attribute_value_mlt":[{"accessrole":"open_date","date":[{"dateType":"Available","dateValue":"2020-05-11"}],"displaytype":"detail","filename":"Gao-2020-Primary orthologs from local sequence.pdf","filesize":[{"value":"3.5 MB"}],"format":"application/pdf","license_note":"Creative Commons Attribution 4.0 International(https://creativecommons.org/licenses/by/4.0/)","licensetype":"license_note","mimetype":"application/pdf","url":{"label":"Gao-2020-Primary orthologs from local sequence","url":"https://oist.repo.nii.ac.jp/record/1471/files/Gao-2020-Primary orthologs from local sequence.pdf"},"version_id":"87859dbe-f3dd-477b-a688-0a77e548f49f"}]},"item_language":{"attribute_name":"言語","attribute_value_mlt":[{"subitem_language":"eng"}]},"item_resource_type":{"attribute_name":"資源タイプ","attribute_value_mlt":[{"resourcetype":"journal article","resourceuri":"http://purl.org/coar/resource_type/c_6501"}]},"item_title":"Primary orthologs from local sequence context","item_titles":{"attribute_name":"タイトル","attribute_value_mlt":[{"subitem_title":"Primary orthologs from local sequence context","subitem_title_language":"en"}]},"item_type_id":"10001","owner":"29","path":["45"],"pubdate":{"attribute_name":"公開日","attribute_value":"2020-05-11"},"publish_date":"2020-05-11","publish_status":"0","recid":"1471","relation_version_is_last":true,"title":["Primary orthologs from local sequence context"],"weko_creator_id":"29","weko_shared_id":29},"updated":"2023-06-26T11:49:44.674230+00:00"}