Search bioRxivSearch

Biology subjects

Francesco Strozzi

Publications and source records attributed to Francesco Strozzi.

2 recordsLinked to original sources

Elucidating the genetic basis of an oligogenic birth defect using whole genome sequence data in a non-model organism, Bubalus bubalis

Recent strong selection for dairy traits in water buffalo has been associated with higher levels of inbreeding, leading to an increase in the prevalence of genetic diseases such as transverse hemimelia (TH), a congenital developmental abnormality characterized by the absence of a variable distal portion of the hindlimbs. The limited genomic resources available for water buffalo, in conjunction with an unconfirmed inheritance pattern, required an original approach to identify genetic variants associated with this disease. The genomes of 4 bilaterally affected cases, 7 unilaterally affected cases, and 14 controls were sequenced. Variant calling identified 19.8 million high confidence single nucleotide polymorphisms (SNPs) and 2.8 million insertions/deletions (INDELs). A concordance analysis of SNPs and INDELs requiring all unilateral and bilateral cases and none of the controls to be homozygous for the same allele, revealed two genes, WNT7A and SMARCA4, known to play a role in embryonic hindlimb development. Additionally, SNP alleles in NOTCH1 and RARB were homozygous exclusively in the bilaterally affected cases, suggesting an oligogenic mode of inheritance. Homozygosity mapping by whole genome de novo assembly was then used to identify large contigs representing regions of homozygosity in the cases. This also supported an oligogenic mode of inheritance; implicating 13 genes involved in aberrant hindlimb development in the bilateral cases and 11 in the unilateral cases. A genome-wide association study (GWAS) predicted additional modifier genes. Results from these analyses suggest that mutations in SMARCA4 and WNT7A are required for expression of TH, while several other loci including NOTCH1 act as modifiers and increase the severity of the disease phenotype. Although our data show that the inheritance of TH is complex, we predict that homozygous variants in WNT7A and SMARCA4 are necessary for the expression of TH and selection against these variants and avoidance of carrier-to-carrier matings should eradicate TH.\n\nAuthor SummaryGenetic diseases often occur and are spread through small populations under strong selection where rates of inbreeding can be significant. The use of a limited number of water buffalo males via artificial insemination for genetic improvement of milk and milk composition has increased the frequency of the genetic disease, transverse hemimelia (TH). Transverse hemimelia affected calves are normally developed except for malformation of one or both hindlimbs or both hindlimbs and one or both forelimbs. Little is known about the inheritance pattern of TH. We discovered genetic variants present in cases where both hindlimbs and one forelimb were affected, cases were both hindlimbs were affected, cases where only one hindlimb was affected, and in non-affected water buffalo that predict TH to be inherited as an oligogenic disease with two driver loci necessary for disease expression and several additional modifier genes that are responsible for the severity of the disease phenotype. We predict that selection against mutations in the two major loci and the avoidance of mating animals that are heterozygous for these mutations will eliminate TH from water buffalo.

Genomics

FALDO: A semantic standard for describing the location of nucleotide and protein feature annotation.

Background Nucleotide and protein sequence feature annotations are essential to understand biology on the genomic, transcriptomic, and proteomic level. Using Semantic Web technologies to query biological annotations, there was no standard that described this potentially complex location information as subject-predicate-object triples.\n\nDescription We have developed an ontology, the Feature Annotation Location Description Ontology (FALDO), to describe the positions of annotated features on linear and circular sequences. FALDO can be used to describe nucleotide features in sequence records, protein annotations, and glycan binding sites, among other features in coordinate systems of the aforementioned \"omics\" areas. Using the same data format to represent sequence positions that are independent of file formats allows us to integrate sequence data from multiple sources and data types. The genome browser JBrowse is used to demonstrate accessing multiple SPARQL endpoints to display genomic feature annotations, as well as protein annotations from UniProt mapped to genomic locations.\n\nConclusions Our ontology allows users to uniformly describe - and potentially merge - sequence annotations from multiple sources. Data sources using FALDO can prospectively be retrieved using federalised SPARQL queries against public SPARQL endpoints and/or local private triple stores.

Bioinformatics