MASARYK UNIVERSITY FACULTY OF MEDICINE DEPARTMENT OF BIOLOGY F U N C T I O N A L S I G N I F I C A N C E O F M U T A T I O N S I N T H E G E N E S A S S O C I A T E D W I T H T H E P R I M A R Y I M M U N O D E F I C I E N C Y D E V E L O P M E N T PhD Thesis in the Field of Medical Biology Supervisor: Author: doc. MUDr. Tomas Freiberger, PhD. Mgr. Lucie Grodecka Brno, 2016 Abstrakt Úvod Genetické varianty mohou být neškodné i škodlivé, přičemž ty škodlivé často podmiňují závažná lidská onemocnění včetně primárních imunodeficiencí (PID). Rozpoznání škodlivých variant není vždy jednoduchou záležitostí, protože řada z nich neovlivňuje přímo funkční produkt genu (například kódující sekvenci proteinu), ale spíše nepřímo narušuje regulaci exprese. Nej častějším typem takových „regulačních" defektů vedoucích k onemocněním člověka jsou poruchy sestřihu pre-mRNA. I tyto aberace však zůstávají často nerozpoznané, zřejmě proto, že regulátory sestřihu mívají velmi degenerované sekvence vyskytující se po celé délce genu. I v případě, že jsou varianty vedoucí k poruchám sestřihu rozpoznány, posouzení jejich vlivu na finální podobu sestřihu a tedy na funkci produktu genu není triviální. Tato práce zkoumá prevalenci mutací ovlivňujících sestřih pre-mRNA u genů spojených se vznikem PID a u několika dalších genů. Dále prozkoumává možnosti rozpoznání těchto mutací a odhadu jejich vlivu na výslednou podobu sestřihu. V neposlední řadě jsou zde navrženy nové mechanizmy aberantního sestřihu. Metody Pomocí in silico predikčních nástrojů bylo analyzováno celkem 111 různých genových variant. 64 z nich pak bylo zkoumáno pomocí in vitro metod s využitím sestřihové minigenové analýzy a RT-PCR z RNA pacientů (s ohledem na dostupnost tohoto materiálu). Výsledky Prevalence sestřihových aberací v genech spjatých s rozvojem PID se ukázala být přibližně 22 %. In silico predikční nástroje přinesly spolehlivé výsledky ohledně poruch míst sestřihu, avšak predikce poruch přídatných sestřihových regulačních elementů se projevily jako málo věrohodné. Přesto se i v této druhé kategorii našly tři nástroje, které vykázaly slibnou efektivitu: konkrétně ESRseq, EX-SKIP a Hexplorer. Kromě toho dva nástroje, konkrétně NNSplice a MaxEnt, ukázaly jistou schopnost rozlišit kryptická místa sestřihu aktivovaná po poruše autentických míst od těch inaktivních. Dále byly určeny specifické hraniční hodnoty, které pomáhají rozlišit varianty v první pozici exonu ovlivňující sestřih od variant tichých. Srovnání minigenové analýzy a RT-PCR z RNA pacientů ukázalo vysokou míru spolehlivosti minigenové analýzy v určování, zda určitá varianta ovlivní sestřih pre-mRNA. Tím může tato metoda pomoci při stanovování patogenního potenciálu nových genetických variant. V neposlední řadě tato práce prokazuje jeden nový mechanizmus aktivace pseudoexonů a předkládá nový možný mechanismus poruchy rozpoznání akceptorového místa sestřihu. Závěr Tyto výsledky pomáhají objasnit význam sestřihových mutací u genů spjatých s rozvojem PID a přispívají k porozumění mechanismům sestřihových aberací. Zde zahrnutá evaluace různých in vitro a in silico metod může být přínosná jak pro genetickou diagnostiku, tak pro vývoj dalších predikčních nástrojů. Klíčová slova sestřih pre-mRNA, primární imunodefícience, varianty nejasného významu, aberace sestřihu, in silico predikce, minigenová analýza Abstract Background Gene variants can be both innoxious and deleterious, the later ones often underlying serious human diseases including primary immunodeficiencies (PIDs). The recognition of the deleterious variants is not always straightforward, as they may not affect the gene function directly (e.g. its protein coding potential), but rather indirectly by disturbing its expression regulation. The most prevalent 'regulatory' defect leading to human disorders is pre-mRNA splicing aberration. Yet even the splicing-affecting variants can be many times unrecognised, as the splicing regulators are highly degenerate and dispersed all over the whole gene. Even when recognised, neither the assessment of the aberrant splicing consequences and its impact on gene function is a simple task. Therefore, this work inspects the prevalence of splicing-affecting mutations in the genes related to the PIDs development as well as in several other genes, explores the ways to recognize such mutations and to assess their impact on splicing pattern, and proposes some novel mechanisms leading to aberrant splicing. Methods One hundred eleven individual genes variants have been analysed by various in silico predictors. Sixty-four of them were later analysed in vitro using splicing minigene assay and RT-PCR of the patients'-derived RNA, when available. Results The prevalence of splicing aberrations in PID-related genes turned out to be approximately 22 %. The in silico prediction tools gave reliable results for splice site affection, while were much less reliable for the changes of splicing regulatory elements. Yet three algorithms, ESRseq, EX-SKTP and Hexplorer showed promising results in the later category. Further, NNSplice and MaxEnt predictors indicated capability of assessing which cryptic splice site would be activated after authentic splice site disruption. In addition, we have proposed specific cut-off values for discerning splicing-affecting from non-affecting variants at the exons first position. Next, comparison of the minigene analyses and the RT-PCR from patients' derived RNA indicated the former method as a useful tool helping assess the pathogenic potential of an inspected variant. Finally, regarding the mechanisms of aberrant splicing, the results showed a novel way of multiple pseudoexon activation and indicated possible novel mechanism of acceptor splice site recognition defect. Conclusions These findings help to elucidate the significance of splicing mutations in the PID-related genes and contribute to the understanding mechanisms of splicing aberration. Further, the evaluation of various in silico and in vitro approaches may come useful for both the diagnosticians and the developers of the in silico prediction tools. Key words pre-mRNA splicing, primary immunodeficiencies, variants of unknown significance, splicing aberration, in silico prediction, minigene assay Declaration of authorship: I hereby declare that I have worked on this thesis entitled "Functional significance of mutations in the genes responsible for the development of primary immunodeficiencies" independently under the supervision of doc. MUDr. Tomas Freiberger, PhD. In addition, Mgr. Kamila Reblova, PhD. and Emanuele Buratti, PhD., gave me valuable advice and comments to the whole text or its parts. Acknowledgement I am sincerely grateful to my supervisor, doc. MUDr. Tomas Freiberger, PhD. for his patient guidance and kind and friendly support for my long own-way-finding journey of PhD studies. I would like to express my special gratefulness to Dr. Emanuele Buratti who always found a while to support me with an advice and encouragement during various projects. My thanks also belong to Mgr. Kamila Reblova, PhD. who gave me valuable recommendations for this text. Naturally, many thanks belong to all my collaborators who have greatly contributed on the process of preparing publications which enabled this thesis to be completed. In addition, I am grateful to my mentor doc. Lucy Vojtova, PhD. who guided me through the last period of the PhD studies and encouraged me to finish the text. Next, my thanks belong also to MUDr Simona Markova who provided me with English correction of several parts of the thesis. Finally and importantly, I would like to express sincere gratitude to all my family, especially to my partner Michal Kajan and to my and his parents, who did their best to help me combine the maternity with finishing the PhD. And I send many kisses to my daughter Michaela who showed both courage and patience letting her mother spend lot of time in 'the other world'. Table of contents 1. Introduction 10 1.1. Splicing 11 1.1.1. Alternative splicing 14 1.1.2. Splicing regulation 16 1.1.2.1. Gs-acting elements 16 1.1.2.2. Tram-acting factors 18 1.1.2.3. Context dependency 18 1.1.2.4. RNA secondary structure 19 1.1.2.5. Signalling cascades 20 1.1.2.6. Tissue specificity 20 1.1.2.7. Rate of transcriptional elongation 21 1.1.3. Splicing aberrations 21 1.1.3.1. Splicing aberration due to defects of trans-factors 22 1.1.3.2. Splicing aberration due to toxic RNA 22 1.1.3.3. Splicing aberration due to disruption of cz's-acting elements 23 1.1.3.4. Effects of splicing disruption 25 1.1.4. Splicing in diagnostics settings 26 1.1.4.1. In silico predictions 26 1.1.4.2. In vitro analyses 35 1.1.4.1. Clinical significance of splicing aberrations 38 1.2. Primary immunodeficiencies 41 1.2.1. Splicing errors inPIDs 41 1.2.2. Splicing-targeted therapies 42 2. Aims of the thesis 45 3. Results and discussion 46 3.1. Variants in the exons first nucleotides: predictions on splicing affection 46 3.2. Splicing-affecting variants in selected PID-related genes 62 3.3. Splice site competition leads to introduction of multiple pseudoexons in the IDS gene.. .86 3.4. Variant c.4211-32_-13del in the FBN1 gene leads to major exon skipping 99 3.5. Variant C.2440-6OG in the CDH1 gene does not markedly affect splicing 106 4. Conclusions I l l 5. References 112 6. Abbreviations 130 7. List of figures 132 8. List of tables 133 9. List of supplements 134 10. List of publications 135 11. Summary of the thesis main findings 136 1. Introduction One of the most awesome biological processes in the living organism is the precisely orchestrated gene expression. Each cell has to finely tune expression of thousands of genes to form its transcriptome and proteome with respect to the cell type, developmental stage, as well as to variable environmental conditions. Yet the process of individual gene expression is as complex as the metamorphosis of larvae into the butterfly. Multiple steps are highly interconnected with each other: transcription, capping, splicing, 3' end cleavage and polyadenylation, RNA editing, nuclear export, translation and posttranslational modifications. Naturally, most of these steps constitute opportunities for regulation as well as for defects that may lead to the development of human disorders. Therefore, the situation of genetic diagnosticians when finding a novel gene variant is not simple at all. They always have to take into account that the variant might influence multiple steps of gene expression. This task is almost never straightforward, since we are only recently getting aware of the complex regulatory mechanisms related to these events, many of which are not yet fully understood. At the same time, we are just entering the genomic era, which enables us to detect thousands of novel variants in a few days. All these reasons highlight the necessity to better understand the whole process of gene expression regulation, to be able to properly assess variants pathogenicity. This work provides an insight into the process of pre-mRNA splicing with focus on the gene variants that lead to its aberration, especially those found in the genes related to the primary immunodeficiency (PID) development. In fact, splicing-affecting mutations are now believed to form a significant portion of defects that underlie the development of many inherited diseases as well as of cancer. Still, it is not always simple to forecast variants effect on splicing, even though many in silico prediction tools exist. Many individual splicing-affecting mutations have been described even in the PID-related genes. On the other hand, as the PIDs symptoms are often explicit and diagnosis is therefore straightforward, there is usually not a big stress on verifying functional significance of the detected mutations. Yet, variants that need to be functionally verified occur occasionally in every gene, most frequently being suspected for splicing affection. Here, we have inspected in total 111 individual gene variants (64 of them in vitro), mostly found in the PID-related genes, for their effect on splicing. Using these data, we have 10 evaluated the reliability of several in silico prediction tools for discerning the splicing-affecting variants from those with no effect on splicing, as well as for the assessing the aberrant splicing pattern. Further, we have proposed possible novel mechanisms of acceptor splice site recognition. Finally, we have described a new mechanism that led to pseudoexon activation. 1.1. Splicing One of the most important features of gene expression that distinguishes prokaryotic from eukaryotic organisms is the existence of RNA splicing in the latter ones. The seminal finding that the genetic information can be separated in discontinuous segments of DNA which are joined together during the process of mature mRNA formation was first described in late 1970s. Since then, remarkable effort has been devoted to understand this process and its regulation, which complexity has led the researchers to tentatively define a "splicing code". Such a set of rules shall once enable to predict splicing pattern of any primary transcript.1,2 During the process of splicing, the borders of intervening sequences (or so called introns) are recognized, cleaved and exons are then ligated together, all of which is catalysed by a large molecular complex termed spliceosome1 In the pre-mRNA, the exons contain sequence to be later translated into amino acid sequence during the process of translation, while the introns mostly do not code for protein sequence but can contain other genes such as miRNA, snoRNA and tRNA genes and even independent transcription units for both coding and non-coding genes (reviewed by Hube and Francastel).4 However, even exons that do not code for protein sequence exist, either in untranslated regions of the genes or in the genes not coding for proteins, as splicing has been demonstrated for long non-coding RNAs (IncRNAs) as well.5 , 6 2 The spliceosome is composed of five small nuclear ribonucleoprotein (snRNP) complexes: U l , U2, U4, U5 and U6; and an array of non-snRNP protein factors.7 To ensure precise excision of intronic sequences, a multitude of RNA-RNA, RNA-protein and protein-protein interactions successively arise and vanish in the cycle of spliceosome assembly and catalysis. While the biochemical splicing reaction occurs in just two transesterfication steps, the 1 Special groups of self-splicing introns that are not dependent on the spliceosome exist, e.g. in pre-rRNA genes of Tetrahymena and in some pre-mRNAs and pre-rRNA of chloroplast and mitochondrial genomes.3 Further details are out of the scope of this thesis. 2 Even pre-tRNA genes contain introns, but their splicing proceeds through a mechanism very different from all the other types of introns, not involving the transesterification (explained below).3 Further details are out of the scope of this thesis. 11 spliceosome assembly involves dynamic rearrangement through several distinct ribonucleoprotein complexes (Figure l).8 , 9 B (activated) p r P 2 " A T P (catalytically activated) Figure 1. Pre-mRNA splicing by the U2-type spliceosome. Adapted from Will and Luhrman.9 A) Schematic representation of the two-step mechanism of pre-mRNA splicing. Boxes and solid lines represent the exons (El, E2) and the intron, respectively. In the first step, the 2' hydroxyl of the branch point (indicated by the letter A, as it is most frequently adenosine) attacks the 5' terminal phosphodiester bond of the intron. In the second step, the 3' hydroxyl of the upstream exon (El) attacks the 3' terminal phosphodiester bond on the intron, thus joining of the two exons together.9 B) Assembly and disassembly pathway of the U2-dependent spliceosome. Exons and introns are indicated as in the A part. Circles depict individual snRNPs and their successive interactions. The DExH/D-box RNA ATPases/helicases Prp5, Sub2/UAP56, Prp28, Brr2, Prp2, Prpl6, Prp22 and Prp43, or the GTPase Snull4 facilitate conformational changes occurring during individual assembly stages. 12 Already in the early complexes (E- and E-complex), both intron ends are occupied by their specific trans-factors, in particular U l snRNP at the 5' splice site and the SF1 (or later U2 auxiliary factor; U2AF) at the branch site or 3' splice site, respectively (Figure 2). An issue that markedly influences the splicing results is whether the 5' splice site and 3'splice siteoccupying factors interact with each other over exon or over intron in these early complexes. It was shown that spliceosomal assembly is more effective when splice sites are reasonably close to each other, while it becomes ineffective when their distance exceeds 250-300 nt.10 Hand in hand with this notion, initial splice site-pair recognition seems to occur mostly crossexon (through so called "exon-definition"), at least for human genes with typical short-exons, long-introns architecture.8 AGBranch point Polypyrimidine tract E' complex Rearrangement to C complex and catalysis Exon definition Figure 2. Early complexes of the spliceosome assembly. Adapted from Chen and Manley.8 The possible transition steps between exon and intron definition are shown by arrows, the transition step supported by later finding of Schneider et al.n is encircled. 5'ss and 3'ss indicate donor and acceptor splice sites, respectively. Circles depict individual trans-acting factors. Cylinders and solid lines represent the exons and the introns, respectively. 13 Besides this "major spliceosome", a minor spliceosome form has been identified. This low-abundant spliceosome differs from the major one mainly in being composed of four specific snRNPs (Ull, U12, U4atac and U6atac) that are functionally analogous to the U l , U2, U4 and U6 snRNPs, respectively. The minor spliceosome is responsible for splicing of a specific class of rare introns, so called U12-type introns (as opposed to U2-type introns recognized by a major spliceosome). It is estimated that circa 700 - 800 human genes possess some introns of this type. These genes are involved e.g. in D N A replication and repair, transcription, RNA processing and translation. On the other hand, U-12 introns are scarce in the genes related to basic energy metabolism and biosynthetic pathways.12 1.1.1. Alternative splicing Considering the energy that the cells have to expend into the replication of intronic D N A sequences, their transcription and subsequent splicing by the large cellular apparatus, a question that logically arises is: what are the advantages of preserving such a complicated process? In fact, splicing may provide many evolutionary advantages, such as protection from mutations and transposons, expanded genome coding capacity together with proteome expansion and another level of gene expression regulation.13 Indeed, splicing of most human genes is highly regulated in a process known as alternative splicing. This term basically describes a situation in which at least one splice site competes for its recognition and joining into mature RNA with another splice site, thus producing distinct RNAs from an identical primary transcript (Figure 3).1 3 In the case of pre-mRNA splicing, resulting transcripts may differ in their protein coding potential, leading to the production of diverse protein isoforms. Furthermore, alternative splicing can impact the fate of mRNA transcripts: when introducing a premature stop codon, it destines the RNA for degradation through nonsense-mediated decay (NMD)1 ; when altering untranslated regions, it can affect the presence of elements regulating transcript stability, translation efficiency and localization.14 In addition, alternative splicing affects the biogenesis of non-protein coding RNAs such as miRNAs and lncRNAs.5 1 5 Thus, alternative splicing plays a central role in genomic diversity, tissue specificity and ontogenesis. In accordance with its evolutionary importance, the extent of alternative splicing scales with the complexity of an organism. Recent studies assess that between 95 and 100 % 3 nonsense-mediated decay: mRNA surveillance pathway that selectively degrades transcripts that contain premature termination codons (Gardner, 2010) 14 of human genes undergo alternative splicing.8 1 6 Not surprisingly, the most complex splicing patterns use to be described in human brain, one of the most functionally specific tissue.8 p r e - m R N A Figure 3. Modes of alternative splicing with its protein coding consequences. Inspired by Faustino and Cooper.17 A) alternative promoter usage. Protein coding potential is affected according to the location of the translation initiation codon. B) cassette exon (see note 8). C) and D) alternative 5' and 3' splice sites, respectively. E) intron retention. F) mutually exclusive exons. G) and H) alternative termination exons: competition between upstream poly(A) site and downstream 3' splice site in G; competition between 5' splice site and poly(A) site within upstream terminal exon.17 15 1.1.2. Splicing regulation The selection of particular positions to cut and religate into the mature RNA proceeds through multiple interactions of czs-acting sequence elements with their cognate trans-acting spliceosomal and non-spliceosomal factors.18 Yet many other characteristics than just the particular elements and their binding partners markedly impact these interactions, e.g. sequence and chromatin context, RNA secondary structure, and of course the availability and functional state of interacting proteins.18 "20 1.1.2.1. C/s-acting elements Among the cz's-elements, the essential sequences recognized by the core splicing machinery are indicated as "splicing signals". These include: donor splice site (5'ss),4 - branch point siteand acceptor splice site (3'ss)- (Figure 4).1 3 The splice sites are the most conserved among all splicing elements, yet there are only two almost 100 % conserved positions at each splice site, while the requirements for conservation of other nucleotides are much lower.21 "24 The pseudo splice sites, i.e. similar sequences that resemble real splice sites but are never used under normal conditions, can outnumber the real, so called authentic splice sites by an order of magnitude.17,25 Therefore, additional regulatory sequences are necessary to discern between the authentic and pseudo splice sites, as well as to specify the conditions for alternative splicing. These so called "splicing regulatory elements" (SREs) bind splicing activators or repressors and influence (help or impede) the spliceosomal recognition and/or usage of proximate splice sites (reviewed in Cartegni et al.26 ). 4 donor splice site: conserved sequence at the 3' end of exons and the 5' end of introns; typically recognized by U l snRNA 5 branch site: conserved sequence circa 25 nt upstream from the acceptor splice site; typically recognized by SF1 and later by U2 snRNA 6 acceptor splice site: conserved sequence at the 3' end of introns and the 5' end of exons; typically recognized by U2AF heterodimer 16 5' Splice Site Branch Site 3' Splice Site c C 5' Exon v Intron 3' Exon Figure 4. Consensual sequences of splice sites and the branch site. Adapted from Padgett.27 The size of each letter represents the frequency of a base at the particular position. Note that only the sequences of the U2-dependent introns (recognized by the major spliceosome) are shown. SREs are involved in both constitutive and alternative splicing. Compared to splicing signals, their cognate sequences are much more degenerate.28 SREs were originally categorized according to their localization and effect on splicing as exonic or intronic splicing enhancers (ESEs and ISEs) and silencers (ESSs and ISSs; Figure 5).1 7 This rather simplistic view was later reassessed, as many elements acting as ESEs function as silencers when found in introns, and vice versa. Further studies have shown that position and context dependence is paramount to the function of most of the splicing elements, both the splicing signals and SREs. In addition, most regulation contexts emerge from overlapping RNA elements.13 Interestingly, some of the overlapping elements may have both silencing and enhancing functions, as originally implicated for a composite exonic regulatory element of splicing (CERES).2 9 Figure 5. Basic categorization of splicing regulatory elements (SREs). Boxes and solid lines represent the exons and introns, respectively. Enhancers and silencers are denoted according to their exonic or intronic locations as ESE, ISE, ESS and ESS (see paragraph above the Figure). Ovals represent trans-acting splicing factors, either the spliceosomal ones (shown in orange) or those binding the SREs (shown in other colours). Activating and inhibitory interactions are shown in black. U l and U2 stand for U l and U2 snRNAs, respectively. SR denotes SR-proteins, hn indicates hnRNPs. 17 1.1.2.2. Trans-acting factors Two groups of protein factors, SR-proteins and heterogeneous nuclear ribonucleoproteins (hnRNPs) are known as classic trans-acting splicing regulators. Still, there are many other factors that play role in constitutive as well as alternative splicing regulation. SR-proteins, named according to their typical arginine- and serin-rich domains (RS-domains) were originally described as general splicing factors, only to be later recognized as sequence-specific splicing regulators.8 These proteins play important roles in many other aspects of RNA metabolism, such as mRNA nuclear export, NMD and translation.8 Their exonic cognate sequences mostly act as splicing enhancers. Upon binding to an exon, SR proteins can recruit U l snRNP to the 5'ss and U2AF to the 3'ss through protein-protein interactions (in early steps of spliceosome assembly), thus promoting exon inclusion. Of note, SR proteins can even operate in an opposite manner under certain circumstances.13,30 Antagonistic function to SR proteins is generally attributed to hnRNPs that have been commonly involved in interactions with splicing silencing elements, ESSs and ISSs.2 8 , 3 0 The term hnRNPs was initially chosen to describe the group of proteins that associated with heterogeneous nuclear RNA or pre-mRNA.3 1 Not surprisingly, these proteins have been later demonstrated to play role in a variety of biological processes other than splicing regulation, e.g. in telomere biogenesis, polyadenylation, translation, RNA editing and mRNA stability. 1 3 , 3 1 Several distinct mechanisms have been shown in hnRNPs-mediated splicing silencing, e.g. block of interaction between U l snRNP and U2 snRNP, direct inhibition of the snRNP binding, prevention of SR proteins binding to pre-mRNA, looping out of an exon, etc.31,32 Yet some factors of this large protein family were demonstrated as splicing activators through binding to ISEs.1 3 , 3 2 1.1.2.3. Context dependency As mentioned above, an important feature of splicing elements is their high context dependency. In fact, function of all splicing elements depends on their localization and sequence context. For example, the strength2 of and the distance between two splice sites influence each other's recognition, as their binding factors (Ul snRNP and U2 snRNP) interact in the process of exon definition (see above in the chapter 1.1 'Splicing').13 Another 7 splice site strength: the quality of a splice site as a degree of its resemblance to the consensus and, consequently, the likelihood of being effectively recognized by the spliceosome 18 classical illustration of context dependency is that some sequence elements behave as splicing enhancers when located in exons, while they repress splicing from intronic positions.33 To complicate the matter a little bit more, some SREs act adversely from various intronic positions. E.g. NOVA binding elements activated alternative exon inclusion when bound to the downstream intron, but repressed its inclusion when bound to the upstream intron.34 Finally, different effect of the same element was demonstrated when the element was placed elsewhere in the same exon.35 Not all of these events have been successfully explained. It is clear, however, that all activities of splicing regulators communicate with the function of core spliceosome, at first for the splice sites definition and then for modulation of selective pairing of the defined splice sites. It is conceivable that such coordination can function differently when a regulator interacts from a position downstream of a 5'ss or upstream of a 3'ss. In addition, some regulators were shown to involve cooperative binding to their cognate elements. In many other instances, competition of different factors to bind overlapping elements has been demonstrated. Moreover, due to high degeneracy of the SREs, the same motif can serve as a binding site for several different regulators.13 Complicating the matter a little bit more, the elements' recognition is influenced also by its accessibility, i.e. state of chromatin and RNA secondary structure.13 '36,37 It has been shown that some chromatin modifiers interact with splicing factors and that certain splicing events (e.g. some cassette exons-) have specific chromatin modification patterns. Presently, it is hypothesised that certain chromatin signatures might favour recruitment of specific factors, or possibly vice versa, certain splicing factors might recruit specific chromatin modifiers.37 Therefore, the role of chromatin state in the splicing regulation is highly probable. 1.1.2.4. RNA secondary structure RNA secondary structure may affect splicing basically in two ways. Firstly, secondary structures may prevent binding of general splicing factors to their cognate elements. Secondly, the secondary structure formation can markedly change the distance between regulatory elements, thus influencing the efficiency of splice site usage. These rules hold true both for splice sites and SREs. As an example, the "looping out mechanism" is well known for hnRNP A l protein which binds to the intronic sequences at each site of one exon of its own pre-mRNA. The two protein molecules then dimerize, so the exonic sequence between them 8 cassette exon: an alternative exon that may be either included into mRNA or skipped (see Figure 2) 19 forms a loop which then helps exclude that exon from the line of available exons.38 Naturally, not only the binding of some proteins to RNA depends on its secondary structure, but the proteins bound to RNA affect its structural folding as well. As the folding of nascent RNA occurs simultaneously with gene transcription, the formation of secondary structure is tightly connected to the transcriptional rate. Taken together, the selection constraints set on exonic sequences are triple: i) preservation of coding sequence; ii) preservation of SREs and iii) preservation of appropriate structural context for splicing regulatory motives.39 1.1.2.5. Signalling cascades Another way of alternative splicing regulation lays at the level of the post-translational modifications of splicing regulators. Well known is the regulators' phosphorylation that influences their ability to interact with RNA and with other protein factors, or affects their subcellular localization. In some cases, differently phosphorylated protein may once act as splicing activator and then as splicing repressor.8 As many splicing regulators affect alternative splicing of a great variety of genes, their modification seems to be an ideal means for the cellular stress reaction.18 Indeed, changes in alternative splicing have been demonstrated as a response for osmotic stress, DNA damage, heat shock, etc.13 1.1.2.6. Tissue specificity As the alternative splicing is an important driver of tissue specificity, there are several ways the cells use to achieve the diversity.18 Firstly, each cell type has its unique repertoire of transacting splicing regulators, such as SR-proteins and hnRNPs.8 Given the highly complex cooperation of all regulatory components, it is perceivable that proper alternative splicing regulation depends also on the stoichiometry and interactions of positive and negative regulators (including core spliceosomal proteins) associated with particular RNA. Secondly, it was shown that splicing machinery prefers either high exon inclusion or high exon exclusion. Such a threshold kinetics provides for marked effects on alternative splicing pattern even upon subtle changes in the regulators stoichiometry.8,40 In addition, splicing factors often regulate their own expression as well as expression of other splicing regulators.13 One of the most interesting mechanisms is "unproductive splicing of SRprotein genes". These genes contain so called 'poison exons': coding sequences that contain a premature termination codon (PTC) located more than 50 nt upstream from the last exon-exon junction, thus directing the RNA to degradation in the process of NMD. SR-proteins direct 20 alternative splicing of their own poison exons, thus regulating their own expression.41 1.1.2.7. Rate of transcriptional elongation Another level of splicing regulation occurs due to its interconnection with transcription. In fact, splicing of most exons occurs cotranscriptionally, only minority of exons is spliced post transcription.42 Therefore, introns tend to be spliced sequentially, upstream introns before downstream introns. As alternative splice sites compete for their recognition, slow elongation should favour the usage of upstream splice sites, resulting in higher inclusion of alternative cassette exons. With the same logic, fast elongation should reduce the competitive advantage of upstream splice sites and promote exon skipping. At the same time, the elongation rate could influence the competition between splicing enhancers and silencers.43 This so called "window for opportunity" model was indeed demonstrated for a fraction of cassette exons. However, a portion of exons showed an opposite effect: higher inclusion upon faster transcription. Interestingly, the two types of splicing responses were connected to distinct sequence features of the responding exons. Most interestingly, a major group of exons showed the same direction of the response on both slower and faster than wild type elongation rate, what suggests that these exons demand "just right" rate of transcription for proper level of alternative splicing.43 This cannot be explained by the "window of opportunity" model. Fong et al.4 3 suggested that splicing of these exons is markedly influenced by two partially independent parameters: nascent RNA folding and protein binding to splicing signals and SREs. This and many other findings emphasise that the RNA splicing is a part of much wider gene expression regulating network. In fact, all the processes of mRNA biosynthesis including transcription, capping, splicing, polyadenylation and nuclear export are tightly interconnected, with a multitude of their regulatory factors playing role in more than one of these sub- processes.32 Therefore, the number of parameters that affect splicing regulation is huge. 1.1.3. Splicing aberrations As every biological process, even the splicing is prone to errors, including both the random anomalies and the systemic splicing aberrancies. While the former ones arise from natural inaccuracy of the splicing process, the latter ones generally result from a change of genomic DNA.2 8 , 4 4 Naturally, the systemic splicing aberrations constitute a significant mechanism that leads to many human diseases. There are basically three major ways through which a genetic 21 mutation affects the process of splicing: i) defect in cis, ii) defect of a trans-acting factor and iii) aberrant mRNA transcript-driven toxicity.2 8 , 4 5 The most common splicing aberrations occur as effects in cis. They can affect virtually any sequence element including splice sites, branch site, SREs. In addition, they can even impair these elements in an indirect way, influencing their accessibility by changing RNA secondary structure. Splicing defects in cis have been described in the majority of Mendelian disorders, including many PIDs.2 8 , 4 5 As this category of splicing aberrations is paramount for this theses, this subject will be further expanded below. 1.1.3.1. Splicing aberration due to defects of trans-factors Much rarer than cz's-elements disruptions are the defects of trans-acting splicing factors. Several of them have been described in a variety of hereditary disorders such as retinitis pigmentosa,46 ^18 amyotrophic lateral sclerosis,49,50 spinal muscular atrophy (SMA),5 1 and dilated cardiomyopathy.52 Despite these aberrations affect splicing of a whole array of transcripts, they often lead to an impairment of one specific tissue or a cell type. Possibly, some tissues may have higher need for a specific splicing factor as has been proposed for SMA.5 3 For example, it seems that some of the transcripts that are mis-spliced in SMA are vital for the motor neuron function.53 In addition, other roles of the mutated factors than in splicing regulation may be crucial for the disease phenotype, such as snRNP cytoplasmnucleus trafficking or mRNA transport in case of SMA and ALS, respectively.53,54 At least one defect of a trans-acting splicing factor related to immunodeficiency has been identified so far. In 2015, Merico et al.5 5 described mutations in RNU4ATAC gene which led to the development of Roifman Syndrome characterized with antibody deficiency and spondyloepiphyseal chondro-osseous dysplasia, retinal dystrophy, poor pre- and postnatal growth, cognitive delay and facial dysmorphism. The mutated gene codes for U4atac snRNA which is essential part of minor spliceosome. Resulting differential expression of XRCC5 in patients with Roifman Syndrome could be responsible for their defective immunity, since the product of this gene takes part in double-strand break repairs, a process important also during T-cell and B-cell receptor V(D)J recombination.55 1.1.3.2. Splicing aberration due to toxic RNA As a third category, effect of toxic RNAs, also known as spliceopathy, has been demonstrated in several degenerative disorders that are characterized by repeat expansion, e.g. in myotonic 22 dystrophy or spinocerebellar ataxia.45,56 "58 Basically, when a repetitive sequence forms a binding motif for a trans- acting factor, the expanded binding sites recruit and sequester that particular factor (mostly RNA-binding protein), leading to its depletion in the cell and major splicing deregulation.59 "61 To the best of my knowledge, such an RNA pathology has not been related to any immunodeficiency yet. 1.1.3.3. Splicing aberration due to disruption of czs-acting elements In theory, mutation of any cz's-acting splicing element may result in defective splicing of the mutated pre-mRNA, such as exon skipping, aberrant splice site2 —— activation, intron retention or affection of the balance of individual alternatively spliced isoforms.28 In particular, splice site disruption and creation is the most frequently described cause of aberrant splicing. Another important category comprises disruptions and creations of SREs, followed by somewhat rarer affection of RNA secondary structure. Disruptions of splice sites frequently lead to aberrant splicing. Particularly mutations of the conserved dinucleotides (GU/AG in 99 % of introns) result in splicing disruption in almost 100 % of cases.62,63 In the rest of splice site sequences, the probability of a mutation harmful for splicing roughly corresponds to position conservation.21,64,65 Besides the disruptive mutations, creation of a new splice site consensus or "strengthening" of an existing cryptic splice site (increasing the match to a consensus) have been many times found as a splicing spoiler. Such a sequence use to be recognized by the spliceosome instead of the original (authentic) splice site, in all or part of the transcripts.64 Next, splicing aberrations due to disruption of a branch site have been rarely demonstrated.66 Not all the mutations of this element are pathogenic: the spliceosome can often use cryptic branch site located nearby or even the mutated nucleotide. Further, it was shown that some introns possess more than one functional branch site.67 SRE-affecting mutations resulting in aberrant splicing have been described many times.2 6 , 6 8 , 6 9 Particularly, ESE disruptions mostly lead to exon skipping, either in a portion or in all of the mutated transcripts.28,70 "72 Similar effects were reported for ESS-creating mutations.73 "77 Interestingly, a change C.840OT between SMN1 and SMN2 paralogous genes (a mutation 9 aberrant splice sites: cryptic or de novo splice sites (Vofechovsky06) 10 cryptic splice sites: sequences that resemble splice sites but are used only when their authentic counterpart is disrupted 11 de novo splice sites: sequences in which a variant newly forms or increases splice site consensus 23 from evolutionary point of view) evidently induces both ESE disruption and ESS creation at the same time.7 8 , 7 9 This substitution results in predominant exon skipping in most of SMN2 transcripts, what hampers this gene from preserving normal SMN function in SMA patients suffering from SMN1 deficiency.80 '81 Considering the degeneracy of SREs, the concurrent ESE disruption and ESS creation may be a common mechanism of SRE affection, as supported by other reports.82,83 In addition, ESE creation or ESS disruption might also affect splicing, either by disrupting the balance of alternatively spliced isoforms or by activating a cryptic splice site in the vicinity of affected regulatory element. Indeed, ESS disruption in alternative exon 4 of CD45 gene was shown to affect alternative splicing of its gene, leading to a predisposition to multiple sclerosis.84 (Lynch and Weiss, 2001). The creation of "exonic" enhancers has been implicated in pseudoexon activation, which is addressed in more detail below. An analogical situation concerns intronic enhancers and silencers: both their creation and disruption can seriously affect splicing. More specifically, deletion of an ISE in PLP1 gene affected normal balance of alternatively spliced transcripts, leading to a mild form of Pelizaeus-Merzbacher disease.85,86 On the other hand, creation of an ISE resulted in an activation of a pseudoexon— in the MTRR gene, resulting in homocystinuria.87 Similarly, ISS disruption in CHRNA1 gene increased inclusion of alternative exon P3A leading to massive production of a non-functional protein, causing congenital myasthenic syndrome.88 As an opposite example, an intronic substitution that emerged during evolution of SMN genes created an ISS in SMN2 intron 7 which binds hnRNPAl with high affinity and probably contributes to the major exon 7 skipping in the SMN2 transcripts.89 In addition, very striking can be the consequences of CERES mutations, where different alterations of the same element lead to opposite splicing outcomes.29,90,91 The SRE disruptions prove the high context dependence of the splicing regulation. Although positive and negative SREs are abundant in all pre-mRNA, not all of these elements seem to be necessary for exon recognition.86,92 "94 Frequently being redundant, only a portion of them could be recognized by /ram-acting factors in a given gene and given tissue.86 An illustrative example is a SNP in exon 5 of MCAD gene that inactivates herein localized ESS: while having no effect on its own, this polymorphism alleviates the need of nearby enhancers. Hence, otherwise harmful mutations leading to major exon 5 skipping were proven harmless 12 pseudoexon: intronic sequence that resembles exon but is neglected by the spliceosome under normal conditions; may be activated by a mutation 24 in the polymorphic sequence context.95 Therefore, it is still extremely challenging to recognize whether a mutation does or does not affect splicing through an SRE change.72,96 Nonetheless, it emerges that some exons are more susceptible to SRE disruptions than others. Typical signs for these susceptible exons are: i) beyond-average length (either short or particularly long, a factor probably connected to splice site recognition through exon-definition process) ii) weak splice sites and iii) alternative splicing.1 0 , 9 7 "9 9 In addition, other related factors that influence the basic level of exon recognition and might thus influence exons' susceptibility to SRE-affecting variants are RNA secondary structure, nucleosome density and transcription rate.99 Last but not least, a mutation that affects RNA secondary structure can disrupt normal splicing process. An illustrative example is the case of 5'ss of exon 10 of the tan gene. Sequence of this splice site normally forms a hairpin in the pre-mRNA thus making it hardly accessible to the spliceosome. Mutations that disturb this secondary structure facilitate U l snRNA binding to the 5'ss and increase exon 10 inclusion into the mRNA, leading to parkinsonism.1 0 0 1 0 1 1.1.3.4. Effects of splicing disruption In general, the most prevalent effect of splice site disruption is exon skipping, followed in its frequency by cryptic splice site usage.102 1 0 3 This choice is mainly influenced by the availability of cryptic splice site in the vicinity of the disrupted site and by the intrinsic properties of the splicing regulators in the mutated sequence context.104 1 0 5 In other words, exons that were skipped after splice site mutation were shorter than average exons, had significantly weaker 3'ss, higher density of predicted ESSs and lower of predicted ESEs.1 0 6 Interestingly, a mutation of a czs-regulatory element (e.g. SRE or splice site) may result in skipping of more than one exons in the mRNA, both upstream and downstream from the mutated exon.1 0 7 "1 0 9 In addition, intron retention or affection of the balance of individual alternatively spliced isoforms are less frequently described consequences of splicing disruption, compared to exon skipping or activation of aberrant splice sites.2 8 1 0 4 A specific result of splicing-affecting mutations is an inclusion of a pseudoexon into the mRNA. Pseudoexons are intronic sequences of cca 50 - 200 nt length flanked by apparently viable consensus splice sites at both ends. Despite being very abundant in human genome, their inclusion into mature mRNA is rather scarce.30 Still, a mutation can induce pseudoexon inclusion using several mechanisms that may or may not directly affect the pseudoexonic 25 sequence. These are: i) creation or strengthening of a splice site at one pseudoexon boundary,110111 ii) creation of a branch site,112 iii) disruption or creation of a SRE in the pseudoexonic sequence or in its close vicinity,2 9 , 8 7 1 1 3 iv) disruption of an authentic splice site of the neighbouring exon,1 1 4 - 1 1 7 v) ESE disruption in the neighbouring exon, that led to the bypass of its authentic splice site (proposed by Stucki et al.1 1 8 ) and finally vi) gross genomic rearrangements that bring together splice sites that would normally be far from each other.119 In addition, also these events can be markedly influenced by the pre-mRNA secondary structure, as shown in the cases of CFTR and ATM pseudoexon inclusion.65 1.1.4. Splicing in diagnostics settings As the splicing aberrations often underlie serious genetic diseases (both the inborn ones and the cancer), there is a considerable effort to recognize them, understand their impact on gene expression and in some cases even to correct them. Although the detection of a DNA change has been a trivial task since the development of next generation sequencing techniques, recognition of a splicing affecting event is still a challenge, mainly due to the high degeneracy and context dependence of the splicing code.120 Another difficulty then arises from the uncertainly in the minimal extent of aberrant splicing that can lead to a particular disease. In other words, aberrant splicing may not always unambiguously lead to the development of a disorder, e.g. when the resulting transcript causes only minor change in the encoded protein or when residual amount of wild type transcript is sufficient to cover the biological function.121 ~ 1 2 3 To address these issues, many in silico and in vitro approaches have been developed, as well as several elaborated guidelines that help genetic diagnosticians to opt for an optimal clinical management. 1.1.4.1. In silico predictions An easily applicable approach to assess mutations' effects on splicing are the in silico predictions. There exist a number of individual prediction tools and even several engines that combine multiple tools in one interface. Many of them are online, freely accessible and their usage is generally simple. Basically, to assess an effect of a mutation on splicing we can use predictions of splicing signals (splice sites, branch point, polypyrimidine tract), of SREs and of RNA secondary structure. In addition, more complex approaches (e.g. those based on splicing code) has emerged recently as well. 26 Splice site predictions Among the splicing predictions, the highest reliability is usually reported for splice site-predicting tools (Table l ) . 6 3 9 6 1 2 4 This is not surprising when one realizes that the splice sites are the most conserved of all the splicing elements. Original computational assessments of splice site quality were based on position weight matrices derived from statistical analysis of several thousand known splice sites.125 Later, such matrices were improved using much more primary sequence data21 or accounting for the interdependencies of adjacent or even non-adjacent positions, such as in the first-order Markov model and the maximum-entropy model.65 In addition, many other approaches have been used, such as decision tree method (program GENSCAN),1 2 6 machine-learning based on neural network (NNSplice),1 2 7 assessment of the quality of 5'ss binding to the U l snRNA in terms of hydrogen bonds between the two partners (Hbond and #H).6 5 , 1 2 8 Of all this variability, the maximum entropy model-based tools usually show the best performance - both in discriminating the splicing-affecting mutations from the non-affecting ones6 3 , 1 2 9 and when distinguishing the authentic from the aberrant splice sites.64,65 Comparable performance was shown for Splice Site Finder, a position-weight matrix models-based program.63 However, this tool is not freely accessible any more, as it became a part of Alamut software (Interactive Biosoftware). Not surprisingly, the most reliable predictions on splicing affection are usually obtained for the changes of the invariant dinucleotides (GU/AG in U2-type introns).63 In addition, predictions were shown much more reliable when the authentic splice site was strong (i.e. similar to the consensus), whereas the disruptions of weakly defined splice sites were much more difficult to identify.63 Importantly, the predicted score often do not correlate with the splice site usage, because of the effect of local sequence context.25,130 On the other hand, Shepard et al.4 0 showed that the level of alternative exon inclusion correlated with the sum of the splice site scores predicted for the exon. Importantly, the prediction programs give users just the score (or sometimes even the score change). The user has to determine his/her own criteria that distinguish between the predicted pathogenic and harmless variants. Such criteria, mostly cut-off values, were often set arbitrarily in different studies, frequently as 10 % difference between wild type and variant sequence score.96,131 "133 Later, more extensive studies of splicing-affecting mutations enabled computation of receiver-operator characteristics (or ROC curves) which then allowed to 27 determine optimal cut-off values, i.e. optimal compromise between sensitivity and specificity for individual prediction tools.63 1 3 4 However, one has to keep in mind that these values might be specific for their primary data set or for a specific gene. The generalizability to other genes has yet to be tested.63 1 3 5 In addition, some studies recommended to combine several tools to profit from their advantages while overcoming their drawbacks.63,96 Accordingly, Jian et al.1 3 4 showed that their two models that combined several splice site prediction tools outperformed each of the individual tools. Prediction tool Principle Web site Ref. Position weight matrix probability of a nucleotide at the position using statistical comparison with large data sample http://sroogle.tau.ac.il/ 125 MaxEntScan position weight matrix based on maximum entropy model of short sequence distribution, includes dependencies between adjacent as well as non-adjacent positions http://genes.mit.edu/burgelab/rn axent/Xmaxentscan scoreseq.ht n_ http://genes.mit.edu/burgelab/rn axent/Xmaxentscanscoreseqa cc.html 136 GENSCAN probabilistic model based on decision tree method http://genes.mit.edu/GENSCAN html 126 NNSplice neural network (machine learning) http://www.fruitfly.org/seq_tool s/splice.html 127 Hbond analysis of hydrogen bonding patterns between 5'ss and Ul snRNA https://www2.hhu.de/rna/html/h bondscore.php 128 Analyzer Splice Tool counting the number of hydrogen bonds between 5'ss and U l snRNA (among other programs) http://host- ibis2.tau.ac.il/ssat/SpliceSiteFra me.htm 59 HSF position weight matrix http://www.umd.be/HSF3/index. html 131 Table 1. Selected online tools for prediction of splice site changes. Another issue that has yet to be addressed b the community of splicing researches is the prediction of the aberrant splicing pattern that emerges when the splicing is affected. For 28 obvious reasons, the particular outcome of a splicing-affecting mutation is crucial for diagnostic decisions (see chapter 1.1.4.1 'Clinical significance of splicing aberrations' for further explanation). However, to the best of my knowledge, there were only a few pioneer attempts to use in silico prediction tools in this area.1 2 9 1 3 7 Specific acceptor splice site predictions Several branch point prediction tools exist (Table 2), using position weight matrices,1 3 1 1 3 8 complementarity to U2 snRNA1 3 9 '1 4 0 or a support vector machine algorithm.141 Corvelo et al. 1 4 1 showed that the BP predictions are more effective when they take into account other 3'ss properties, such as polypyrimidine tract (PPT)— or AG-exclusion zones,14 probably due to high interdependence between individual elements. However, as there is only a few welldescribed BP mutations, existing evaluation of these predictors is very limited.66 Similar situation applies to PPT predictors: their efficiency has not been extensively tested and the development of new tools is rather oriented at predicting more complex sequence elements (Table 2). Prediction tool Principle Web site Ref. 'branch site predictor — position weight matrix http://sroogle.tau.ac.il/ 138 'branch site predictor' position weight matrix http://www.umd.be/HSF3/index.html 131 'branch site predictor' complementarity to U2 snRNA http://sroogle.tau.ac.il/ 140 SVM-BP finder support vectormachine algorithm http://regulatongenomics.upf.edu/Software/SV M B P / 141 "PPT predictor' position weight matrix http://sroogle.tau.ac.il/ 138 'PPT predictor' algorithm quantifying pyrimidine enrichment http://sroogle.tau.ac.il/ 140 Table 2. Selected tools predicting specific 3'ss features: branch point and PPT. 11 polypyrimidine tract: pyrymidine-rich sequence stretch between the branch point and 5' end of exon, part of 3'ss 14 AG-exclusion zones: regions upstream of 3'ss AG, devoid of other AG dinucleotides 15 prediction tools between the apostrophes do not have any special name 29 Predictions on SREs Following the most striking finding of diagnostics-related splicing research made in the last decade of the 20t h century that virtually any D N A change may affect splicing through disrupting SREs, many attempts to define these elements occurred.26 Some of the systemic approaches then led to the development of SRE-predicting tools (Table 3). As the first, Liu et al.1 4 2 used functional systematic evolution of ligands by exponential enrichment (SELEX) to define enhancer motifs bound by four classical SR-proteins, SRSF1, SRSF2, SRSF5 and SRSF6. Based on these matrices, Cartegni et al.1 4 3 then developed an online tool called ESEfinder. The same SELEX approach was later used to assign binding preferences to other splicing factors, such as hnRNPAl, Tra2beta, 9G8.1 3 1 Fairbrother et al.1 4 4 presented another pioneer approach to the identification of ESEs: Relative Enhancer and Silencer Classification by Unanimous Enrichment (RESCUE). They compared the occurrence of all possible hexanucleotide motifs in exonic and intronic sequences and in exons with weak and with strong splice sites. The rationale behind such experiment is that ESEs are more prevalent in exons than in introns and that exons with weak splice sites are more dependent on the presence of ESEs.1 4 4 The authors made these data easily accessible in the online tool named according to method they have used: RESCUE-ESE.1 4 5 Other scientists later adapted the same principle for ESEs' and other SREs' identification, only using different rationales.35 1 4 6 "1 5 1 Another way of defining functional SREs works through using splicing reporter systems. Wang et al.1 5 2 adopted a systematic minigene analysis to test random decanucleotides for their ESS properties. More specifically, they inserted random sequences into the middle exon of a three-exon minigene. The middle exon was constitutively spliced, unless an ESS was present in the inserted sequence. Later, Ke et al.1 5 3 used analogous approach to assign enhancing or silencing properties to all possible hexanucleotides. Importantly, in order to make provision for the context dependence of the hexamer activities, authors tested all the sequence motifs in five different exonic locations. Unfortunately, no online program based on this approach has yet been made available to the public. Recently, two prediction tools based on splicing code modelling (using Bayesian neural network) occurred. Program AVISPA (Advanced Visualization of Splicing Prediction and Analysis) predicts alternative splicing and tissue specificity of particular exons, as well as it can identify putative regulatory elements. However, it was trained on mouse alternative 30 splicing data, which limits its usage for human genes.154 The second tool, SPANR (Splicingbased Analysis of Variants) is designed for prediction of effects that a nucleotide variant exerts on cassette exon alternative splicing.155 In addition, Piva et al.1 5 6 presented their database of human splicing factors expression data and RNA target motifs named SpliceAid2. Here, users can get predictions on potential SREs which are based strictly on comparison with experimentally proved sequence motifs shown to bind splicing regulatory proteins. In parallel, such a collective approach was adopted also by Raponi et al.1 5 7 who found the best predictions on mutation-induced exon skipping when they combined multiple individual predictions together. Thus developed tools, EX-SKIP and HOTSKIP, count the sum of ESEs and ESSs derived from individual predictions and then calculate the ESS/ESE ratio, either for the specific variant (in EX-SKIP) or for each possible single nucleotide substitution in a selected exon (in HOT-SKIP). An advantage of this tool is that even the individual predictions are digestedly shown on the results page, so the user does not have to approach each program individually. Such an advantage is also held by several online engines - web pages, that include several individual prediction tools at one location, e.g. Sroogle and Human Splicing Finder (Table 4).131158 Prediction tool Principle Web site Ref. ESE-finder SELEX (in vitro selection of ligands) http://krainer01.cshl.edu/cgi- bin/tools/ESE3/esefinder.cgi? process=home 143 RESCUE-ESE statistical comparison of hexamer sequence motifs http://genes.mit.edu/burgelab/ rescue-ese/ 145 PESX statistical comparison of octamer sequence motifs http://cubio.biology.columbia .edu/pesx/pesx/ 147 FAS-ESS minigene analysis of sequence silencing properties http://genes.mit.edu/fas-ess/ 152 SPANR splicing code, machine learning http://tools.genes.toronto.edu/ 155 SpliceAid2 database of splicing factors binding sites www.introni.it/spliceaid.html 156 Hexplorer statistical comparison of hexamer sequence motifs http://nar.oxfordjournals.0rg/c ontent/early/2014/08/21/nar.g ku736/suppl/DCl 151 Table 3. Selected tools for prediction of SREs. 31 Despite the developers of SRE-predicting tools mostly proved their functionality in independent in vitro analyses, additional studies often did not find these predictions efficient.13124 In fact, Královicova and Vořechovský1 0 6 and subsequently Woolfe et al.8 3 showed that these tools generally do recognize motifs with splicing-regulatory properties, as they were able to statistically distinguish sequences with different propensity to activate cryptic splice sites or the splicing-affecting variants from SNPs, respectively. Similarly, Raponi et al.1 5 7 indicated several statistically significant correlations between SREs predictions and exon inclusion. However, difficulties arise when studies test these predictors for resolution of individual splicing-affecting from harmless sequence changes.71 '72 '83 '97 1 5 9 1 6 0 In fact, Chasin1 6 1 reported that at least 75 % of nucleotides in a typical human exon are found within an SRE motif predicted by some of the available tools. In accordance, a study that indicated ESE-finder and RESCUE-ESE predictions as rather successful (able to identify 6 out of 7 ESE-disrupting mutations) was based on a priori proved ESE-disrupting mutations, excluding the possibility to assess programs specificity.162 Other studies indicated SREpredicting tools as less efficient, inconclusive and difficult to interpret.71,97 However, the number of such studies and the quantity of included SRE-affecting sequence variant is still rather limited. Interestingly, two studies have shown some promising achievements of recently developed algorithms in recognition of splicing affecting variants, at least in BRCA2 and MLH1 genes.72 1 6 3 These algorithms were ESRseq scores and Hexplorer.1 5 1 1 5 3 Still, several of the possibly important variables that affect the variants' effects on splicing are not involved in most of the prediction algorithms, e.g. proximity of certain exons (related to exon/intron definition of splice sites), alternative splicing, local RNA structure, nucleosome density and rate of pre-mRNA synthesis.8,99 32 Prediction tool Included programs and other approaches Web site Ref. EX-SKIP PESE and PESS,147 FAS-ESS,152 RESCUE-ESE,1 4 5 EIEs and IIEs,150 NI-ESE and NI-ESS164 http://ex- skip.img.cas.cz/ 157 HOT-SKIP PESE and PESS, FAS-ESS, RESCUE-ESE, EIEs and IIEs, NIESE and NI-ESS http://hot- skip.img.cas.cz/ 157 Sroogle branch site prediction,140 PPT prediction,138140 MaxEnt,136 Position-specific scoring matrix,123 delta-G, ESE-finder,143 RESCUE-ESE, FAS-ESS, and other SREs35 1 4 8 1 4 9 http://sroogle.ta u.ac.il/ 158 Human Splicing Finder (HSF) HSF-specific matrix for splice sites, MaxEnt, HSF-specific branch point calculation, ESE-finder, RESCUE-ESE, PESE and PESS, EIEs and IIEs, FAS-ESS and ESS decamers,152 Exonic splicing regulatory sequences,33 and other HSF- specific matrices for SREs http://www.um d.be/HSF3/inde x.html 131 Table 4. Selected online engines combining multiple splicing-prediction tools. Predictions on RNA secondary structure Concerning secondary structure assessment, there are principally two approaches: one based on the minimum free energy calculations and the other based on base pairing probability.1 6 5 1 6 6 On these principles, several online tools have been developed (Table 5). Originally, they only predicted possible structures of individual sequences. Recently, several programs occurred that specifically assess the secondary structure differences caused by single nucleotide variants (SNVs).1 6 6 In general, secondary structure predictions are more accurate for highly structured transcripts, i.e. those that adopt a few well-defined conformations - often related to a specific regulatory function. Such structures are often formed by inverted repeats, as opposed to average pre-mRNAs that generally adopt an ensemble of diverse conformations.166 1 6 7 Yet, global efficiency in pre-mRNA secondary structure predictions was shown e.g. as these helped to improve predictions of splice sites.168 1 6 9 On the other hand, individual predicted RNA secondary structures were confirmed by in vitro studies in some but not allc a s e s 1 5 7 1 5 8 .1 6 7 .1 7 0 1 7 1 33 Larger evaluation of the applicability of these predictions for assessment of individual variants effect on splicing is missing due to limited amount of in vitro verified data. Prediction tool Principle Web site Ref. Mfold minimum free energy structures http://unafold.rna.alb any.edu/?q=mfold 172 Pfold evolutionary probabilistic model http://www.daimi.au.dk/~compbio/pfold/ 173 RNAfold minimum free energy structures and base pair probabilities (multiple programs at one site) http://ma.tbi.univie.ac.at/cgi- bin/RNAfold.cgi 174 RNAsnp comparison of base pairing probabilities http://rth. dk/resources/rnasnp/ 175 remuRNA minimum free energy structures, with Boltzmann distribution probability https://www.ncbi.nlm.mh.gov/CBBresearch/ Przytycka/index.cgi#remurna 176 SNPfold computation of Pearson Correlation Coefficient between the nucleotide pairing probabilities of the two sequences http://ribosnitch.bio.unc.edu/downloads/snpf old/ 177 Table 5. Selected online tools for prediction of RNA secondary structure and its changes. Fortunately, as the RNA secondary structure changes influence many other processes of gene expression than just splicing (e.g. RNA editing, RNA interference, RNA stability), RNA secondary structure predictions are not limited to splicing-affection. Corley et al.1 6 6 took advantage of a large set of in vivo analysed secondary structures and SNVs inducing local structure changes (so called riboSNithches)16 -. Using this data set, they evaluated the ability of several predictors to discern between riboSNitches and SNVs with no effect on RNA secondary structure. They showed that the programs that were specifically developed for recognition of structure-affecting variants outperform the more general tools. Further, predictors based on base pairing probability matrix performed better than those based on algorithms calculating minimum free energy. The best efficiency was demonstrated for tools 16 riboSNithches: single nucleotide variants that induce a local secondary structure changes in RNA 34 RNAsnp, remuRNAand SNPfold.1 6 6 1 7 5 "1 7 7 Briefly, the RNA secondary structure predictions may be very helpful in particular cases when one seeks for structure changes near a known splicing regulatory element or when such an element is highly suspected. On the other hand, sole structure change does not imply splicing affection and a negative prediction does not mean that no change in RNA secondary structure occurred. 1.1.4.2. In vitro analyses RT-PCR Undoubtedly, the most direct and reliable way to establish an effect of a DNA variant on splicing is to use RT-PCR analysis of patient's endogenous RNA. On the other hand, such an approach is not always applicable for several reasons: firstly, the patient is not always available; second, available samples are not always suitable for analysis (due to tissue-specific gene expression or possible tissue-specific splicing); third, one has to account with possible aberrant transcript degradation by NMD.1 2 0 Indeed, NMD could be blocked in tissue cultures by cultivation of patients' lymphoblasts. However, its time and costs requirements do not predispose this method to become a routine practise for every potential splicing-affecting variant. In addition, one has to be very careful for all the gene variants in order to assign to every transcript its underlying variant.120 1 2 4 Further, when comparing the results to those of healthy controls, one has to have a respect to significant physiological variability between individuals, in particular regarding alternative splicing events.124 RNA sequencing When patient's samples are available, another method of choice is R N A sequencing (RNAseq). With the expansion of next generation sequencing methods, the possibility of using RNAseq to detect splicing aberrations (as well as other alterations of gene expression) become reachable. In fact, this method can capture and quantify plenty of splicing isoforms at once, thus being especially useful for investigation of multiple targets, e.g. panel of genes (cancer-related, PID-related, etc.). So far, there remains a question of sample availability in sufficient quantity to ensure necessary sequencing depth. In addition, straightforward connection of these results to a biological function is impossible.124 35 Hybrid minigene systems Hybrid minigene is a term used for a simplified gene usually cloned into a plasmid vector that enables researchers to dissect and analyse specific regulatory events. For splicing minigene assays, the most frequently used minigene consists of two vector-specific exons with an intron in between that involves a cloning site (Figure 6). Into this site, the researchers clone a genomic fragment of interest, most frequently one inspected exon together with its intronic flanking regions. Then, the constructs are transfected into cultured cells and after 18-48 hours of cultivation, the resulting processed minigene RNAis analysed using RT-PCR.1 6 5 , 1 7 8 intron Ean intron Figure 6. Typical minigene preparation. Most frequently, the analysed exon (Ean) with parts of its neighbouring introns is cloned into the multiple cloning site (MCS) which lies inside intron surrounded by two exons (El and E2) unrelated to the Ean. Other abbreviations: pr. = promoter; R = resistance gene for antibiotics selection; ori = replication origin. 36 Interestingly, several more specific minigenes for capturing splicing-affecting variants were prepared. These include either vectors that are specifically modified for cloning of the first or the last exon of a gene,179 1 8 0 or minigenes that enable deeper inspection of enhancing/silencing properties of a particular sequence.152181 The main advantage of this approach is that it is relatively fast and easy way to identify splicing aberrations and their principals.120 Despite its often questioned reliability at the beginning, recent extensive study showed almost perfect concordance between patients' RNA analysis and minigene assays in the assessment of aberrant splicing.1 2 2 In addition, this method markedly assisted in resolving complex regulatory events, such as mutually exclusive exons or the impact of promoter architecture on splicing regulation.124 Among the disadvantages of this method the one the most emphasized is the limited and unnatural sequence context. As mentioned above, researchers mostly opt for the easiest way of preparing minigenes and clone just one exon with parts of surrounding introns. However, some variants found in a particular exon affect splicing of the neighbouring one or even multiple exons, as mentioned in the section 1.1.3.4 'Effects of splicing disruption'.107 "109 Hand in hand with this, results of minigene analysis sometimes differ quantitatively according to the particular vector used in the study, possibly due to the differences in promoter architecture or pre-mRNA secondary structure.178181 Further, usage of various cell lines can sometimes bring different results, mostly in terms of transcripts relative quantity9 7 1 8 1 At least in one reported case, a variant affected splicing in one cell line while it was even silent in another one.97 Yet other studies showed only minor or no differences between various cell lines.7 1 1 2 2 As follows from what was noted above, results obtained with this approach have to be always taken as "minigene-specific", especially when a variant causes a change in the proportion of splicing products.120 In vitro splicing Another possibility to assess splicing aberrations is the method of in vitro splicing. During its course, in vitro transcribed pre-mRNA is incubated with nuclear extracts and subsequently resolved on denaturing polyacrylamide gels. Advantageously above minigenes, this method enables to capture splicing intermediates and to assess splicing kinetics. On the other hand, its disadvantages lay in the upper sequence length limit (cca 2000 nts) and the even more artificial conditions for splicing compared to minigene assays.120 For these reasons, the in 37 vitro splicing found its usage much more in basic research than in the diagnostics setup.124 RNA-protein interactions analyses For a deeper inspection of an effect of a mutation on splicing, methods for analysis of RNA-protein binding may come useful. Briefly, Cross-Linking and ImmunoPrecipitation (known as CLIP) enables to detect both individual as well as multiple RNA-protein binding events right in the cellular environment. The a priori selection and often even optimization of the method for each analysed protein is necessary. On the contrary, RNA affinity purification assay (called also RNA pull-down assay) takes place in tubes, where RNA is incubated with nuclear extracts. In theory, it allows to compare binding patterns of all nuclear proteins and to detect different binding to variant sequence of a priori unknown protein, subsequently using mass spectrometry analysis or sequential western blotting with multiple antibodies. Other existing methods are either more suitable for analysis of multiple splicing events (microarrays-based techniques), or are devoted more to basic research of RNA-protein interaction kinetics (electromobility shift assay).1 2 4 1 8 2 1.1.4.1. Clinical significance of splicing aberrations Regarding the huge complexity of both the causes and the consequences of splicing aberrations, one can easily envision the challenges that a genetic diagnostician meets when working with novel gene variants, virtually any of which might potentially be a splicing-affecting variant. Naturally, gene variants may affect any other step of gene expression as well, often with hardly predictable consequences (including many changes of protein-coding potential). Therefore, extensive effort has been put into defining classification criteria for assessment of variant pathogenicity, as this issue markedly impact patients' clinical management. Spurdle et al.1 8 3 published a pioneer work specifically aimed on incorporation of an effect of variants on splicing into these guidelines. Their proposals were later re-assessed and modified by Walker et al.1 8 4 Subsequently, Thompson et al.1 2 1 adopted these criteria for classification of variants found in mismatch repair genes, together with general IARC classification system— for cancer-susceptibility genes (suggested by Plön et al.1 8 5 ). All these works define 5 classes of pathogenic probability. Variants are classified as: pathogenic (class 5) or likely pathogenic (class 4); clinical testing and high-risk 17 IARC classification system: International Agency for Research on Cancer classification system that "conveys information about the relevance of a variant to clinical practice including cancer risk assessment and the need for future studies" (Plön et al., 2008) 38 surveillance guidelines are suggested; uncertain (class 3); likely not pathogenic (class 2) and not pathogenic (class 1); should be treated as if "no mutation for this disorder" has been detected Class 2, 3 and 4 variants are advised for additional data acquisition to enable more robust classifications.121 Generally, all the guidelines concord on the criteria defining class 5 which are, for obvious reasons, very conservative. Among splicing-affecting variants, only those leading to aberrant splicing of 100 % of transcripts, with clear affection of protein function (i.e. introducing PTC or in-frame deletion disrupting a functional domain) fall into this category.121184 Moreover, this criterion can be used as a stand-alone rule only when the splicing aberration has been assayed in patient's RNA. If such a result was derived from in vitro assays (such as splicing minigene assay), Thompson et al.1 2 1 allow class 5 ranking on the basis of additional clinical evidence (family co-segregation, number of tumours, etc.). Into class 4 fall the splicing aberrations in the conserved dinucleotides of the splice sites— that are either not tested in vitro111 or which wild type transcript production was not clearly excluded.184 In addition, Thompson et al.1 2 1 proposes to include variants that abrogate mRNA using in vitro assays only (without patient's RNA-derived data), if there is additional clinical evidence for the variants' pathogenicity. It follows that many variants fall into class 3 - uncertain. Among splicing-affecting variants, this holds true for example for those resulting in: in frame deletion not disrupting known functional domains leaky aberrant splicing (variant allele produces wild type transcripts to some extent) upregulation of physiological alternatively spliced transcripts, at least some of which do not clearly disrupt protein function aberrant splicing detected solely in the minigene assay with insufficient other clinical evidence.1 2 1 1 8 4 18 and the last nucleotide of the exon if the variant is G>non-G and the first 6 bases of the intron are not GTRRGT, specifically according to Thompson et al., 2014 39 In the case that no splicing aberration is observed (variant allele produces the same transcripts as found in healthy controls, supposed that NMD is inhibited), it should be classified as class 2 variant and multifactorial likelihood analysis is recommended.121184 Naturally, effects of the variants on other processes then splicing has to be taken into account (mRNA stability, direct impact on protein, etc.). Into class 1 (variants not pathogenic or of little clinical significance) fall among others the variants that either show the frequency higher than 1 % in the control reference group (the disease prevalence and possible founder effects should be always taken into account) or show variant-specific gene product function in in vitro assays and lack of cosegregation with the disease.121184 In fact, most studies suggest patient's RNA examination as the only confirmation of splicing aberration, supporting a careful and cautious approach to minigenes assays.124 1 8 3 1 8 4 On the other hand, van der Klift et al.1 2 2 proposed that minigene assay alone should be sufficient for clinical classification of splicing-affecting variants as likely pathogenic, if they demonstrate complete aberrant and frameshifting effect. Such reclassification could have important consequences for patients, since class 4 and class 5 variants, as indicated above, are recommended to have the same clinical management.122 Another important contribution to the classification of splicing variants came with the publication of Houdayer et al.6 3 that extensively evaluated several in silico splicing prediction tools. Using both in silico and transcript analyses, they defined three classes of variants according to their splicing-affection. Despite these results were intended to refine clinical classification of variants of unknown significance, they still await to be integrated into the classification guidelines. So far, they help prioritize variants for laboratory-based mRNA assays.184 In addition, Walker et al.1 8 4 proposes his own criteria to assess splicing prediction results (and to include them as a prior probability of pathogenicity), while Thompson et al.1 2 1 do not recommend usage of in silico splicing predictions at all. Of note, there is still a lack of reliable and validated tools predicting splicing affection upon mutations of SREs.6 3 , 7 2 , 1 8 6 Future goals Overall, there are many issues that has yet to be addressed to better assess impact of splicing-affecting variants in diagnostics settings. Additionally to what was mentioned above, there is a difficulty with quantification of splicing aberrations. On one hand, it is a practically and technically demanding task to get RNA from NMD-blocked cells of the disease-relevant tissue. On the other hand, individual minimal doses of full length transcript preventing each 40 disease has yet to be determined," at least for the recessively inherited and haploinsufficiency-driven disorders. In parallel, what remains to be elucidated is the impact of changes in alternative transcript levels (mainly of low abundant isoforms), together with functional roles and standard quantities of these isoforms, and their importance in the process of disease pathogenesis.187 Importantly, one has to envisage that all these characteristics may be somewhat variable between individuals, similarly as can be an effect of the same splicing-affecting mutation.29 All the steps towards the clarification of variants functional significance recently become of even greater importance, as the next-generation sequencing methods now produce thousands of novel variants with every genome read that need to be properly classified. 1.2. Primary immunodeficiencies Inborn defects in one or more components of the immune system that result in impaired, disordered or uncontrolled immune responses are collectively called PIDs. These comprise over 260 genetically-defined single-gene traits with X-linked, autosomal recessive and autosomal dominant inheritance.188 While the individual conditions are generally very rare (the least prevalent occur once in 500.000 newborns), the current estimates of total prevalence of PIDs are one in 1200 humans.189 1 9 0 The tell-tale sign of most PIDs is the increased susceptibility to infections, but the disease often include increased susceptibility to tumours and auto-immunity.191 Therefore, the PIDs form an important group of serious, often life-threatening conditions. The high level of PIDs phenotypic and genotypic heterogeneity has made the genetic analysis essential both for its diagnosis and prognosis.192 1.2.1. Splicing errors in PIDs In PIDs-related scientific literature, the problem of variants of unknown significance is not mentioned as often as e.g. in research related to cancer susceptibility genes. When the patient's phenotype is clear and falls well into the defined criteria for the respective condition, and a rare genetic variant is found in a gene of known and well-described relation with the disease development, there is generally lower press to prove functional consequence of the variant on the gene expression. This might be one of the reasons why the splicing-affecting variants, while often found in individual cases, have not yet been systematically inspected in PID-related genes. As indicated above, the high heterogeneity of individual PIDs often makes the finding of a 41 particular genetic variant very helpful.192 A portion of the detected variants are hypomorphic mutations, that disrupt but not totally abolish the function of the affected gene's product. Such mutations often lead to atypical PIDs presentation that can resemble another condition than usually related to the mutated gene, be milder than expected or occur later in life than in typical cases.193 "197 Hypomorphic mutations might be relatively frequent especially among splicing-affecting mutations either due to their leakiness (the tendency to preserve a fraction of normally spliced transcripts)198 "200 or due to residual activity of the aberrant protein product.195,201 Yet, sometimes it is not straightforward to determine which of these too possibilities arose in the particular case.196 There are many peculiarities among splicing-affecting mutations. The consequences of hypomorphic splicing mutations may underlie phenotype differences between affected siblings.1 9 8 , 2 0 2 — Disease causing splicing mutations can be found even in the untranslated regions of the gene, and their detection and recognition as splicing-affecting variants can be equally challenging as that of deep intronic mutations.203 In addition, in some rare cases a splicing-affecting mutation can be suppressed by another mutation that reverses the effect of the previous one, such as in the case of atypical A D A deficiency.195 On the other hand, number of detected splicing-affecting mutations result in complete aberration with typical phenotypic effects. 1.2.2. Splicing-targeted therapies Due to PIDs diversity, even their treatment varies fundamentally between individual conditions. The most common measures (although not quite general) aim at prevention and treatment of infections. More specifically, more than 50 % of the PIDs patients suffer from antibody deficiencies, being treated by replacement of the missing antibodies. More severe defects usually require either the hematopoietic stem cell transplantation or enzyme replacement or correction of the faulty gene.192,204 The gene therapy may take place both at the DNA and the RNA level, where the splicing comes up on the scene again. In fact, researchers can manipulate splicing to restore normal state when splicing is affected by a mutation, to regulate alternative splicing process, to prevent mutated allele expression or to remove causal mutations from RNA.2 0 5 , 2 0 6 Of note, the 19 Of note, there are other causes of phenotype differences between the individuals carrying the same genetic variant, such as effects of disease-modifying genes or more generally genomic context and the environmental factors (Cooper et al., 2013). Still, patient-specific (extent of) splicing affection can be one of the reasons. (Pagani et al., 2003) 42 last two options are not limited for targeting splicing-affecting mutations. In any case, the deep comprehension to the splicing process of the particular gene is essential for any of these options, thus highlighting the importance of splicing research even in the PIDs area. Basically, there are four therapeutic strategies to alter splicing, using: antisense oligonucleotides (ASOs), spliceosomal-mediated R N A /raws-splicing (SMaRT), small molecule compounds and modifications of snRNAs. Probably the most promising approach is the use of ASOs, small stretches of DNA or RNA that may be chemically modified.1 9 2 , 2 0 6 ASOs act by steric blocking, their binding to a target site prevents interaction of trans-acting regulators with the site. They can prevent usage of a splice site (e.g. a cryptic or a de novo splice site), or block binding to an SRE, with the potential of both promoting or preventing splicing at the adjacent splice site.206 In addition, so called bifunctional ASOs have been developed to enhance or prevent exon inclusion, binding the target mRNA by one end and recruiting the splicing effectors by their other end.2 0 6 , 2 0 7 Among PIDs, ASOs have been experimentally tested for use with several splicing-affecting mutations in the ataxia teleangiectasia patients, where other treatment possibilities are limited.2 0 8 , 2 0 9 In addition, since the ASO-based therapy has transient effect, X L A seems to be especially suitable for this therapy. Defective B-cells in the XLA-patients need only a transient restoration of btk expression to overcome developmental block and to generate longliving mature plasma cells.2 1 0 Another way of therapeutic splicing manipulation uses the /raws-splicing, in which the spliceosome combines two pre-mRNA molecules into one final mRNA. This enables replacement of a portion of mRNA molecule by another, which may be artificially delivered (as with SMaRT). In PIDs, the /raws-splicing therapy has been experimentally tested in CD40L-deficient mice2 1 1 and in DNA-PKcs-deficient cell-lines.212 Yet this methodology suffers from several limitations, such as necessity of sophisticated specificity optimization and complications with the delivery of DNA expression vectors to the cells, what prevents it from wider usage.206,213 Alternatively, mutated pre-mRNA may be corrected using small molecule compounds that can directly or indirectly alter splicing. In addition, modified snRNAs (mostly the U l snRNAs) can rescue the cases of splice site mutations or even some cases of other exon-skipping mutations, e.g. using so called Exon Specific U l snRNAs.2 0 6 , 2 1 4 To the best of my knowledge, these approaches have not been tested for PIDs treatment yet. 43 As mentioned above, splicing manipulation can be used as a therapy even in cases when the mutation does not affect splicing. Pre-mRNA bearing a mutation can be corrected by inducing skipping of the mutated and possibly of other exons, if both the reading frame and at least partial function of the protein are preserved. To this end, many approaches including ASOs, SMaRT and small molecule compounds have been successfully used. Moreover, splicing modification can be used to prevent detrimental mutated allele expression. For induction of aberrant splicing leading to NMD-based degradation of the mutated transcripts, both ASOs and small molecule compounds have been proved effective.206 In PIDs, the so called reframing concept has been used e.g. in Nimjegen breakage syndrome therapy, where skipping of two consecutive exons excludes frameshift mutation shared by a majority of patients, from the mRNA. The protein product of this shortened mRNA has been shown associated with much milder phenotype.215 44 2. Aims of the thesis A substantial portion of Mendelian disorders is caused by sequence variants that affect pre-mRNA splicing. However, it is not straightforward to establish the effect of these variants on splicing, as the available methods are limited and of uncertain reliability in case of computational predictions and partly of minigene-based analyses as well. Therefore, pathogenic potential of many possibly splicing-affecting variants is classified as "unknown". This represents a significant burden for genetic diagnostics, since only properly classified variants can help with diagnosis, prognosis and clinical management of the patients and their family members. The work presented in this thesis focuses on the recognition of splicing-affecting sequence variants, in the PID-related genes as well as in other genes. Aims: Systematically explore the splicing-affecting mutations in the PID-related genes in terms of their frequency and localizations that may indicate possibly important and to splicing-affection prone regulatory elements; Evaluate the in silico splicing prediction tools in terms of their reliability and suitability for assessment of variants' impact on splicing; Evaluate the in vitro splicing minigene assay in terms of its reliability to detect splicing-affecting variants and to reflect their effect in the natural genomic and cellular environment; Determine the pathogenic potential and possibly the underlying molecular mechanism of selected gene variants that were found in patients suffering from non-PID Mendelian disorders, in particular: c. 1092T>A variant found in the IDS gene • c.421 l-32_-13del variant found in the FBN1 gene • C.2440-6OG variant found in the CDH1 gene 45 3. Results and discussion 3.1. Variants in the exons first nucleotides: predictions on splicing affection Grodecka L, Lockerova P, Ravcukova B, et al. Exon First Nucleotide Mutations in Splicing: Evaluation of In Silico Prediction Tools. Spilianakis CB, ed. PLoS ONE. 2014;9(2):e89570. doi:10.1371/journal.pone.0089570. In last years, several studies came with detailed evaluation of splice site predictors, eventually leading to elaborated calculation of optimal cut-off values to discern between splicing-affecting and non-affecting variants.63 "65 However, it emerges that some splice site positions might have specific characteristics regarding the requirements for sequence context and may deserve a specific cut-off values for forecasting the effects of their mutation, e.g. the +3 and +4 position of the donor site.1 3 1 , 2 1 6 Similarly, exons first nucleotides (E+ 1 ) form one of the less conserved parts of 3'ss and the estimate of splicing outcomes when mutated is not trivial. In 2011, Fu et al.2 1 7 showed that the best characteristics for prediction of splicing affection in case of E + 1 mutation is the quality of PPT of the affected 3'ss (particularly the length of the longest polypyrimidine stretch; PPS). In our hands, the PPS showed rather poor discriminative power: when used with two of the three E + 1 variants found in our PIE) patients, the splicing-affecting and non-affecting variants were found in the 3'ss with basically the same length of PPS. In silico predictions on these variants using general splice site prediction cut-off values were not very helpful.63 At the same time, there was no specific evaluation of the reliability of splicing prediction for the E + 1 position and therefore no specific cut-off value discerning harmless and harmful E + 1 variants was proposed. Using other database- and literature-extracted E + 1 variants: We have experimentally determined splicing affection of 16 E + 1 variants previously untested for their effect on splicing. We have demonstrated splicing affection in the six of these variants, two of which had been included in the SNP database (dbSNP) without any notion about their potential pathogenicity. We have determined the discriminative sequence characteristics of splicing-affecting and non-affecting E + 1 variants. In addition to PPS, total number of Py in the 25 nt 46 upstream of intron-exon border was the most helpful parameter for assessment of splicing outcome after E + 1 mutation. On the other hand, the overlap of these parameters between the groups of affected and non-affected splicing was too huge to be possibly useful in the diagnostic settings. Hand in hand with that, the particular substitution of the E + 1 nucleotide was highly important for the splicing outcome. The lower tolerance of G>T substitutions as compared to G>A substitutions reflected the evolutionary conservation of the E+ 1 , at least among eutherian mammals. On the other hand, it did not reflect the binding sequence preferences of the 3'ss cognate factor, U2AF heterodimer, as were shown in Wu et al.2 1 8 Our findings were later extended by Lockerova who analysed the same set of minigenes with all the possible E + 1 substituting nucleotides.219 This study showed that in three of the tested exons, while the G>A mutation did not affect splicing, substitutions G>C and G>T did. In the remaining 12 samples, exon splicing was either affected by all the three E + 1 substituents or by none of them. Importantly, in 5 of the 6 exons where all the substituents resulted in aberrant splicing, the extent of aberrant splicing was highest for G>T mutations and lowest for G>A mutations, what is in concordance with our original findings. We have proposed an alternative mechanism of splicing disruption through E + 1 variant in which creation of an ESS or disruption of an ESE overlapping 3'ss might play role. The occasionally observed splicing affection even in the 3'ss with long PPT and SRE predictions fostered this hypothesis. In addition, some cases of ESS and ESE overlapping the E + 1 position have been already described.130,220 I believe that these mechanism deserve further elucidation, hopefully in the next few years. We have evaluated the efficiency of several in silico predictors on splicing outcome of the E + 1 mutations and tentatively defined the discriminative cut-off values. The best performing tool was the MaxEnt, reaching up to 95 % sensitivity and 70 % specificity in the assessment of splicing affection. Rather surprisingly, the specific PPT prediction tools did not perform well in these experiments, underscoring the importance of complex 3'ss recognition through cooperative binding of several individual elements. As other splice site prediction tools performed reasonably well, we took advantage of using them in combination with each other and with further sequence characteristics. With this approach we achieved up to 95 % sensitivity with 90 % specificity. Still, it has to be kept in mind that these cut-off values and their evaluation are based on rather 47 limited amount of original samples. Later findings from Lockerová showed that the combinatorial approach reached only 79.5 % sensitivity and 49.1 % specificity when used on the set of minigenes with all possible E + 1 substitutions.219 In future, we might take advantage of some larger data set of splicing mutation, possibly obtained with a high-throughput approach, to refine the proposed cut-off values. In addition, new discoveries on 3'ss recognition mechanisms might help to stratify diverse categories of 3'ss and to further improve predictions on 3'ss splicing-affecting mutations. 48 OPEN 3 ACCESS Freely available online '0-PLOS 10NE Exon First Nucleotide Mutations in Splicing: Evaluation of In Silico Prediction Tools Lucie Grodecká1 '2 , Pavla Lockerová1 , Barbora Ravčuková1 , Emanuele Buratti3 , Francisco E. Baralle3 , Ladislav Dušek4 , Tomáš Freiberger1 , 2 , 5 * 1 Molecular Genetics Laboratory, Centre for Cardiovascular Surgery and Transplantation, Brno, Czech Republic, 2 Central European Institute of Technology, Masaryk University, Brno, Czech Republic, 3 International Centre for Genetic Engineering and Biotechnology, Trieste, Italy, 4 Institute of Biostatistics and Analyses, Masaryk University, Brno, Czech Republic, 5 Institute of Clinical Immunology and Allergology, St. Anne's University Hospital and Masaryk University, Brno, Czech Republic Abstract Mutations in the first nucleotide of exons (E+ 1 ) mostly affect pre-mRNA splicing when found in AG-dependent 3' splice sites, whereas AG-independent splice sites are more resistant. The AG-dependency, however, may be difficult to assess just from primary sequence data as it depends on the quality of the polypyrimidine tract. For this reason, in silico prediction tools are commonly used to score 3' splice sites. In this study, we have assessed the ability of sequence features and in silico prediction tools to discriminate between the splicing-affecting and non-affecting E+ 1 variants. For this purpose, we newly tested 16 substitutions in vitro and derived other variants from literature. Surprisingly, we found that in the presence of the substituting nucleotide, the quality of the polypyrimidine tract alone was not conclusive about its splicing fate. Rather, it was the identity of the substituting nucleotide that markedly influenced it. Among the computational tools tested, the best performance was achieved using the Maximum Entropy Model and Position-Specific Scoring Matrix. As a result of this study, we have now established preliminary discriminative cut-off values showing sensitivity up to 9 5 % and specificity up to 90%. This is expected to improve our ability to detect splicing-affecting variants in a clinical genetic setting. Citation: Grodeckä L, Lockerovä P, Ravcukovä B, Buratti E, Baralle FE, et al. (2014) Exon First Nucleotide Mutations in Splicing: Evaluation of in Silico Prediction Tools. PLoS ONE 9(2): e89570. doi:10.1371/journal.pone.0089570 Editor: Charalampos Babis Spilianakis, University of Crete, Greece Received September 18, 2013; Accepted January 21, 2014; Published February 21, 2014 Copyright: © 2014 Grodecka et al. This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Funding: This work was supported by an internal grant of Centre for Cardiovascular Surgery and Transplantation, number 201209; and further by the project "CEITEC - Central European Institute of Technology" (CZ.1.05/1.1.00/02.0068) and SuPReMMe (CZ.1.07/2.3.00/20.0045). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Competing Interests: The authors have declared that no competing interests exist. * E-mail: tomas.freiberger@cktch.cz Introduction The generation of functional mRNA from a primary transcript requires the precise removal of introns and the ligation of adjacent exons. Splicing accuracy is ensured by the specific interactions of fran.f-splicing factors with cis-splicing sequence elements, the splice sites, branch site (BS) and additional regulatory elements, splicing enhancers and silencers. These sequences tend to be highly degenerate with the exception of the almost invariant dinucleotides at the exon-intron boundaries. Variations in the highly conserved " A G " dinucleotide in the acceptor splice site (3'ss) generally disrupt splicing of the affected gene. However, changes elsewhere in the 3'ss have consequences that are not as obvious [1,2]. In particular, mutations in the first nucleotide of exons (E+ 1 ) lead to various splicing outcomes that are influenced by the A G dependence of the affected splice site [3]. It has been demonstrated that variations in the E + 1 position in an AG-dependent 3'ss predispose the mRNA to aberrant splicing, whereas this process is not affected when the variation is located in an AG-independent acceptor site [3]. The AG-dependence is a feature of introns that requires an intact A G dinucleotide already for the first step of the splicing reaction, whereas AG-independent introns require the A G only in the second transesterification step [4]. In general, A G dependent introns contain shorter and more degenerate polypyrimidine tracts (PPT) and require the binding of both subunits of U2 snRNP auxiliary factor (U2AF65 and U2AF35) for the first transesterification step to occur, whereas AG-independent introns contain more consensual and longer PPTs located closer to the splice site and do not rely on U2AF35 binding [5,6]. Fu et al. [3] have recently observed that the length of the longest pyrimidine stretch (PPS) in the PPT represents the best differentiating feature for intron AG-dependence; however, it is still difficult to determine the AG-dependence by simply examining the D N A sequence (and especially the potential splicing affection represented by an E + 1 mutation). Currendy, several computational instruments are available that can predict the location of the splice site and its strength in terms of score (resemblance to an ideal consensual sequence) as well as the location and strength of the BS, PPT, and other splicing regulatory elements (SREs). All these factors, when correctly integrated, might make the determination of intron AG-dependence easier. Although the reliability of in silico tools is still limited and differs significantly between various algorithms (results often require experimental confirmation) computational predictions still represent an important starting tool when prioritizing an unclassified variant for functional validation [2]. Originally, prediction tools predicted splice site quality based on nucleotide frequencies of independent positions (e.g. Shapiro and Senepathy matrix) [7]. Following this initial approach, more sophisticated predictive strategies were developed such as machine-learning (used in Neural Network Splice Site Prediction PLOS ONE I www.plosone.org 1 February 2014 | Volume 9 | Issue 2 | e89570 Exon 1st NT Mutations in Splicing Tool, NNSplice) and the maximum entropy model (used in Maximum Entropy Based Scoring Method, MaxEnt) [8,9]. The machine-learning approach recognizes sequence patterns upon training with authentic splice sites sequences [8]. On the other hand, the maximum entropy model approximates short-sequence motifs distribution accounting for nonadjacent as well as adjacent dependencies between positions of the splice site [9]. In parallel to these approaches dedicated at scoring acceptor and donor sites there has also been the appearance of additional strategies aimed at predicting the other basic splicing consensus sequences. For example the existing predictions of BS are based on a position-specific weight matrix, whilst PPT presence and quality is predicted using algorithms that account for length and pyrimidine content [10,11]. Finally, predictions for the presence of SRE elements have also made their appearance. In general, SRE prediction programs were basically built up in two ways. Programs of the first type are based on looking for the presence of SRE motifs derived from statistical analyses of differences in oligomer frequencies between various sequences (such as exons vs. introns, noncoding exons vs. pseudo exons, etc.) [12,13]. On the other hand, the second type of programs is composed by SRE prediction algorithms based on in vitro and in cellulo experimental approaches such as SELEX or splicing reporter system experiments [14,15]. At the moment, no prediction program can be said to fully predict consensus and SRE sequences with 100% accuracy and their reliability seems to differ from system to system. Nonetheless, all studies performed so far agree that using a combination of all these programs represents the best strategy to predict possible splicing outcomes. In this study, we analyzed in vitro the splicing potential of sixteen E + mutations in genes related to the development of human hereditary disorders and tested the efficiency of several freely available in silico instruments to predict the splicing outcome of this mutation type. Being the first report evaluating the performance of in silico tools to predict the outcome of E + mutations we believe that these data could serve as preliminary guidance for genetic diagnostics and might also be helpful to bioinformatic developers to develop future tools. Materials and Methods Ethics Statement The patients provided a written statement of informed consent. The research project was approved by the Ethics Committee of the Centre for Cardiovascular Surgery and Transplantation Brno. Mutations and Minigene Constructs Two sequence variants that affect the first nucleotides in the S7Xgene exons 15 and 18 were detected in patients with X-linked agammaglobulinaemia (see Table SI in File SI). Other herein in vitro analyzed variants were sought in the Ensembl database, H G M D (The Human Gene Mutation Database; http://www. hgmd.cf.ac.uk) and RAPID (Resource of Asian Primary Immunodeficiency Diseases; http://rapid.rcai.riken.jp/RAPID). PCR products encompassing the mutated exons and at least 150 bp of the flanking intronic sequence or the 3' U T R sequence were cleaved using appropriate restriction enzymes and cloned into the pET vector (MoBiTec, Gottingen, Germany). Sequence variants derived from literature were prepared by PCR-based artificial mutagenesis using specific primers harboring the desired mutation and the Pfx polymerase (Life Technologies, Carlsbad, CA). The quality of the inserted sequences was controlled by cycle sequencing on an ABI PRISM 3100-Avant Genetic Analyser (Life Technologies). In silico and other Sequence Analyses The computational predictions were performed using the following web-based resources: the NNSplice [8], the MaxEnt [9], the PSSM [7], ESE-finder [14], RESCUE-ESE [12], Chasin ESE [13], Wang ESS [15], Chasin ESS [13] and BS and PPT predictions according to Kol et al. [10] and Schwartz et al. [11]. These tools, with the exception of NNSplice, were accessed using the Splicing Regulation Online Graphical Engine (Sroogle), and the score percentiles within constitutive exon datasets compiled and presented by this server were also used (http://sroogle.tau.ac. il/) [16]. All score and percentile differences between wild type (wt) and mutant variants were calculated as relative differences, e.g. as follows: (wt percentile - mutant percentile)/wt percentile. In theory, changes of SRE that may overlap the E + 1 position could play an important role in affecting the splicing fate of the exon in which they are located (for further explanation see discussion). However, considering that SRE predictions are rather complex and not all predicted changes would clearly affect splicing, we did not consider all SRE predicted changes (see discussion for more details). In fact, we considered predictions as positive only when the gains of exon splicing silencers (ESS) or losses of exon splicing enhancers (ESE) outnumbered the opposite changes of the respective elements. In addition, we performed SRE predictions twice: in one case we kept the borders of exons and introns when inserted in the Sroogle engine; in the other case we allowed the prediction tool to consider the whole sequence as an exonic one. The reason for this approach was that exonic elements are better explored than intronic ones, and binding of RNA binding proteins to these elements may occur before the exact definition of the intron/exon borders. Since we used the PPS of 3'ss as a marked attribute of AGdependence, we feel it necessary to illustrate the rules we used to define it. As the minimal binding sequence of each U2AF65domain comprises of 4 nts [17,18,19], we considered this length as minimal to define PPS. Of these, we counted the longest uninterrupted stretch of Py that has at least one Py located between the third and the twentieth nucleotide from 3'ss. This choice of length was made because according to ours and previous results we do not expect farther located PY stretches to be involved in splice site definition. Therefore, this kind of localization definition was decided so that short Py stretches farther from the splice site were not counted, but at the same time did not exclude very long stretches. Note that according to herein mentioned rules one PPS (of sequence designated as " F u l l " , gene CAPN3, exon 17; see Table SI in File SI) was counted differently in this work with respect to [3]. Cell Culture and Transfection Procedures HeLa cell line was obtained from the A T C C repository by the E. Baralle's laboratory (LGC Standards S.r.L, Milan, Italy). U-937 cell line was purchased from DSMZ (Leibniz Institute DSMZ German Collection of Microorganisms and Cell Cultures, Braunschweig, Germany). HeLa and U-937 cell lines were maintained in RPMI 1640 medium (Sigma-Aldrich, Prague, Czech Republic) supplemented with 10% fetal calf serum (Zoo Servis, Huntirov, Czech Republic), 2 m M L-glutamine, 100 U / m l penicillin and 0.1 mg/ml streptomycin (Sigma-Aldrich). The HeLa cells (1x10 cells per transfection experiment) were seeded into a 24-well plate one day prior to transfection. U-937 cells (3x10'' per transfection experiment) were seeded into a 24-well plate immediately prior to transfection. Both cell lines were PLOS ONE I www.plosone.org 2 February 2014 | Volume 9 | Issue 2 | e89570 Exon 1st NT Mutations in Splicing transfected using 1.2 |il transfection reagent (XtremeGene 9 for HeLa cells and XtremeGene HP for U-937 cells; Roche Applied Science, Prague, Czech Republic) and 400 ng plasmid DNA per transfection experiment. The R N A was extracted 24 hours posttransfection. RNA Extraction Total RNA from HeLa and U-937 cells was purified using an RNeasy Plus mini kit (Qiagen, Hilden, Germany). The quality and quantity of the R N A was determined by spectrophotometry (Thermo Scientific, Wilmington, DE) and its integrity was checked on a 1.5% agarose gel. Total R N A from peripheral blood stabilized in RNA/afer solution (Life Technologies) was extracted using a RiboPure-Blood Kit (Life Technologies) according to the manufacturer's instructions. RT-PCR RNA (10% of the total yield) extracted from the transfected cells was subjected to reverse transcription using random hexamers and a Transcriptor First Strand cDNA Synthesis Kit (Roche Applied Science, Prague, Czech Republic). The synthesized cDNA (5% of the total sample) was used as a template for PCR using recombinant Taq polymerase (Thermo Scientific, Erembodegem, Belgium) and the following pET vector specific primers, forward: 5'-CAGCACCTTTGTGGTTCTCA-3', reverse: 5'-AGTGCCAAGGTCTGAAGGTC-3'. RT-PCR on RNA extracted from patients' whole blood was performed using BTK gene specific primers (available upon request) and SuperScript One-Step RTPCR with Platinum Taq (Life Technologies, Carlsbad, CA). The resulting fragments were visualized on 2% agarose gels. For quantification, the amplicons were resolved on a 4% denaturing polyacrylamide gel containing 7 M urea in 0.5 times Tris-acetate buffer and GeneTools software (Syngene, Cambridge, UK) was employed. Statistics A median estimate supplied with a min-max range was used as robust summary statistics. The statistical significance of the differences between the computationally predicted scores for AG-dependent and AG-independent 3'ss was evaluated using the nonparametric Mann-Whitney test. The correlation between exon skipping and the predicted scores was assessed using Spearman's rank correlation coefficient. The comparison of mutation frequency to T and other nucleotides, and of prediction of SRE changes in the AG-dependent and AG-independent 3'ss were both accomplished using Fisher's exact test. The specificity and sensitivity estimates for the selected prediction tools were supplied with 95% confidence limits, calculated using Vassar College's online Sensitivity/Specificity Calculator (http://faculty.vassar.edu/ lowry/clin 1 .html). Statistical analysis was performed using STATISTICA software (version 10, StatSoft, Tulsa, OK). A value a — 0.05 was used as limit of statistical significance in all performed analyses. Results Initially, we detected two variants that affected the first nucleotide of exons 15 and 18 in the BTK gene in our patients suffering from X-linked agammaglobulinaemia (Figure 1A, Table SI in File SI). Using in vitro analyses, both splicing minigene assay (Figure IB) and RT-PCR assay on RNA extracted from patients' whole blood (Figure 1C) demonstrated that one of these variants (mutation c.l350G>T in exon 15, designated as "013") clearly affected splicing, whereas the other variant (mutation c.l751G>A at the beginning of exon 18, designated as "HK08") did not. According to the results of Fu et al. [3], the 8 nucleotides long PPS of the mutated splice sites were just below the length range to score as AG-independent sites (that were shown to have PPS from 9 to 16 pyrimidines long). Even the results of the computer predictions were not very decisive (shown in the Tables S2 and S3 in File SI). In fact, all the predicted scores evaluating the overall 3'ss quality decreased when the E + 1 mutations were introduced, with a greater decline observed for the 013 variant. By contrast, the PPT predictions did not conform to these data and depicted the 013 PPT to be of the same or even higher strength in comparison with the PPT of the HK08 sequence. Because of this difference, we decided to evaluate the overall reliability of splicing prediction tools for E + 1 mutations, using our mutations and a series of database-derived and previously analyzed E + 1 variants [3]. To minimize interfering factors, we decided to preserve the design of the experiments made by Fu et al. [3] and to include only G + 1 substitutions into our analyses. Using these criteria, we initially compiled three sets of mutations: the first set (named "test set") of 25 sequence variants was used for primary analysis of in silico tools performance; the second set (named "Fu-mut set") comprised of 30 non-natural sequence variants in which PPT was artificially mutated in [3]; and the third set (named "borderline set") consisted of 5 natural databasederived sequence variants selected in such a manner that their PPS were of borderline quality (with regards to length and/or distance from 3'ss) to score as either AG-dependent or AG-independent (Figure 2 and SI and Table SI in File SI). These two latter sets were used to evaluate our results driven from primary analysis. More specifically, the "test set" of mutations was obtained by combining all the 14 naturally occurring sequence variants analyzed by Fu et al. [3], 2 BTK mutations described above and 9 other variants derived from mutation databases. These 9 additional mutations were selected in the genes responsible for the development of primary immunodeficiencies (PID), 4 of them being presumably splicing-affecting and 5 others non-affecting, based on the length of their PPS (Table SI in File SI). All the novel variants were then analyzed using splicing minigene assay the same way as described for the BTK mutations. As shown in Fig. 3A, not all the supposedly splicing-affecting variants were capable of affecting the splicing process and vice versa. In particular, in accordance with our expectations BTK exon 10 (designated as PID1) and FAS exon 8 (PID4) variants did affect splicing whilst CD55 exon 9 (PID8), IL17RA exon 12 (PID9) and NBN exon 13 (PID5) did not. On the other hand, BTK exon 17 (PID2) together with SerpinGl exon 4 (PID3) and ^TMexon 10 (PID7) together with ATM exon 32 (PID6) did not conform with our assumptions, the latter ones showing that they affect splicing. In total, therefore, 13 sequences of the test set were shown to affect the splicing process in a minigene context, whilst the remaining 12 did not affect this process (see Table SI in File SI for details). In the predicted scores, percentiles and their differences between the wt and mutant sequences of the test set, we then searched for a distinction between the sequences vulnerable to the G + 1 mutations and those that were not. The results of the predictions are summarized in Figure 4 and Tables 1 and 2 and depicted in detail in Tables S2, S3 and S4 in File SI). Our results demonstrated that the maximal discriminative power (the maximal statistical significance estimated using the Mann-Whitney test) was achieved using the MaxEnt tool, followed by the PSSM instrument. More specifically, the MaxEnt tool provided a good resolution of intron AG-dependence using mutant sequence scores and percentiles and, in particular, differences between wt and mutant sequence scores and percentiles. Note that all the predicted PLOS ONE I www.plosone.org 3 February 2014 | Volume 9 | Issue 2 | e89570 Exon 1st NT Mutations in Splicing A 013 agagccaaaaggagaagactagtccttgcctttcctgtagGAATCTTTCCCATGA ... G>T HK08 aagactqcaaaccaatttaataatttttttcaccttctagGGGTTTTGATGTGGG ... G>A Figure 1. Analysis of the two E 1 mutations of the BTKgene, 013 (c.1482G>T) and HK08 (c.1883G>A). (A) Schematic sequences of mutated acceptor splice sites. Introns are shown in lower-case, and exons are shown in capital letters. Mutated nucleotides are bold and underlined. The PPS are singly underlined and other polypyrimidine stretches are dashed underlined. (B) RT-PCR of minigenes transfected into HeLa cells. cDNA bands originating from 013 mutated minigene are numbered as follows: 1) cryptic 3'ss utilization 81 nt upstream of the authentic splice site (the aberrant exon starts at c.1350-81G), 2) normally spliced RNA, 3) skipping of mutated exon. (C) RT-PCR from RNA extracted from patients' blood. P = patient's sample, HC = healthy control sample. The 013 cDNA bands are numbered as in (B). doi:10.1371/journal.pone.0089570.g001 percentiles are derived from the Sroogle engine and pertain to the ranking of the predicted values in the dataset of more than fifty thousand constitutive exons [16]. The PSSM showed very similar outcomes, although with lower statistical significance than the MaxEnt tool. Even the NNSplice predictions discriminated between the splicing-affecting and non-affecting G + mutations using the predicted mutant sequence score and the score difference (Table 1). Finally, a somewhat lower discriminative power than that of the tools predicting overall strength of 3'ss was observed for the distance between the BS predicted according to Kol et al. and the acceptor site and for the PPT length predicted according to Kol et al. [10]. On the other hand, other parameters predicted for the BS did not exhibit statistically significant differences, and a number of the PPT predicted parameters (e.g. the scores and percentiles derived from both matrices) failed to discriminate between the two types of introns (Table 2). In parallel to these considerations, in the test set of mutations we also assessed the distribution of additional parameters that were not predicted computationally (see Table 2 for full details). In particular, the observation of statistically significant differences in the length of the PPS and the number of pyrimidines in the 25 nt sequence upstream of the acceptor site between the splicingaffecting and non-affecting G + 1 mutations was in agreement with the results reported by Fu et al. [3] (Table 2). In accordance with the same report, we did not observe a correlation between these parameters or the computationally predicted 3'ss parameters and the level of exon skipping in the set of 13 splicing-affecting G + mutations, with one single exception: the distance of BS from the 3'ss predicted according to Schwartz was correlated with the mutant sequence exon skipping, reaching a value of R=— 0.59 (p = 0.04). With regards to other factors that can also influence exon definition such as 5' splice site (5'ss) strength and exon length we did not observe any difference in these parameters between the splicing-affecting and non-affecting mutations, even after combining the 5'ss score with the predicted 3'ss strength (data not shown). However, a striking difference was evident when comparing the identity of mutant nucleotides: the majority of mutations to T led to aberrant splicing, whereas the majority of mutations to A did not result in any change (see Table SI in File SI). When we tested this hypothesis using Fisher's exact test, comparing the number of T versus non-T mutations in both groups of splicing events (affected and not affected) we reached significance at p = 0.0002. These data conform well to the evolutionary conservation of the exon first nucleotide. Analysis of the herein examined sequences of the test set and the borderline set (30 sequences in total) in 36 eutherian mammal species (using Ensembl database) showed strong G + conservation, and only changes to A were detected in the E + positions of some species. Moreover, a change to T was PLOS ONE I www.plosone.org 4 February 2014 | Volume 9 | Issue 2 | e89570 Exon 1st NT Mutations in Splicing Seq. Fu06 Fu8 Fu7 Fu3 Fu2 PID2 PID3 PID4 013 HK08 Fu9 F u l l PID1 Ful FulO PID6 Fu5 Fu4 Ful 2 PID9 PID5 Ful 3 Ful 4 PID8 PID7 Excn narr.e PKHD1 CLCN2 C0L1A2 EYA1 FECH BTK SERPING1 FAS BTK BTK LAMA2 CAPN3 BTK GH1 CAPN3 ATM HEXA LPL NEU1 IL 17RA NBN C0L6A2 C0L1A2 CD55 ATM Exon number 25 19 37 10 9 17 4 8 15 18 24 17 10 3 10 32 13 5 2 12 13 8 23 9 10 Exo:i length 12 3 bp 74 bp 108 bp 90 bp 165 bp 119 bp 135 bp 25 bp 217 bp 158 bp 144 bp 12 bp 55 bp 12 0 bp 169 bp 133 bp 105 bp 234 bp 193 bp 42 bp 273 bp 27 bp 99 bp 21 bp 372 bp I n t r o n .GGTTCCATGACAGAATTTACCAGAAATGTAACCATCTCAG .GATGACATTGCCCCCTGGGGCCTTATGTTTGGGAACACAG .GGGTCGGAATACCAGAGCTGTAAC-GTTTATTTCCAACAG •GAAGTATACATGTTCTTCACCTGTCATATTCTTATTTTAG .AGGCATAGTCCACTTACGCATTCGTCTATCTTTCGCATAG .CCTCCAAATCCTAATGCAACAAGTCCTGAATCCCTTGCAG .TCATCCTGCAAGTATCTTTCATCTCTGCCCTTTGTTGCAG .TTTTATTTGTCTTTCTCTGCTTCCATTTTTTGCTTTCTAG ,AGAGCCAAAAGGAGAAGACTAGTCCTTGCCTTTCCTGTAG ,AAGACTGCAAACCAATTIAATAATTTTTTTCACCTTCTAG .TTGCCGTKATAAACTCTGAGGGTCTCTTGTCTTTCCTCAG .ATTCACATCTGAAGCATCTTCCTTTCTGTTTCTTCTCAAG .TATGACCAGGAGCCACTCAAGCAGCACTCTCCCTTCACAG .GGTCICCAGCGTAGACCITGGTGGGCGGTCCTTCTCCIAG • AACTCTGTGACCCCAAATTGGTCTTCATCCTCTCTCTAAG .ACCAATACGT GT TAAAAGCAAGTTACATTTTCTCTTTTAG .AATACAGGGCCCAATCTGGCACATGCCCCTTTTCCTCCAG •AAATITACAAATCTGTGITCCTGCTTTTTTCCCTTTTAAG .ACAAGTTTGTCTTTGTTGACCCTTCCTCCTCCCATGACAG .TCCCAAAGAACCACTCCAGTGTATTTCTTTTCCTTTCCAG ,AACTTGATTATTTACCTIGCATTCITTTCTTTTTCTACAG .AGAGC TCCTCACTAATGCCCCTCTCTCCTCCTGCCCCCAG .ATCCTCCTCCTCTATCTGTTTTTTTTTTTTTTTTGAATAG .AATTITTGTTGTTAATCCTTTTTTICCCCTTCGTCTGIAG .ATGGAATAGTTTTCAAATTATCCTTTTTTTTTTTTTTTAG -IIExon Change GTCTCTGATGAAÄAC. GGAGCGCAGASTCGG. GGTGCTGCTGGTCAA. GATCCACCCACTTCA. GTTGGTCCAATGCCC. GTAIGTCCTGGATGA. GGGCTGGGGAGAACA. GAAACAGTGGCAATA. GAATCTTTCCCATGA. GGGTTTTGATGTGGG. GTGAATGTGGAAGGC. GT7CCCAAAGAGATG. GTGGTATTCCAAÄCA. GAAGAAGCCTATATC. GCTCCATGGTTCCTA. GAAATTAACCATTTT . GCCCAGAGCAGGGGC. GCCTCGATCCAGCTG. GTGCAGCCGCTGGTG. GGCCTGGAAGTGAAA. GGATTTGAGTGAAAG . GGCGTTCCTGGCTTC. GGCCCTCCTGGTAGT. GTACTACCCGTCTTC. GCTACAGAT TGCAAC. B Seq. Exon name BOR3 BRCA2 BOR5 ATP7A BOR2 BRCA2 BOR1 COL1A1 BOR4 NF1 Exon Exon number length I n t r o n 24 16 12 11 24 303 bp 186 bp 9 6 bp 108 bp 84 bp -II- Exon . . .GGAATCTCCATATGTTGAATT7TTGTTT7GTTTTCTGTAG GTTTCAGATGAAATT. . . . TGGAGGATCATCAAGTCATTGTATCTTAATTTTTTTACAG GTAAAGGTAGTGGTA. . . .TTATTTGCCTTAAAAACATATATGAAATATTTCTTTTTAG GAGAACCCTCAATCA . . . . G T T C T A A T G G C C C T T C C T T G T C T T C T T C A T C T C T C T C C A G GGTGCTCGAGGATTG. . . . A A C A T T G T T T G C T G T T T C T C T T T T C T C C A C C A T T C T A T A G GAATAAGATGGTAGA. Change G>A G>A G>T G>T G>T PLOS ONE I www.plosone.org 5 February 2014 | Volume 9 | Issue 2 | e89570 Exon 1st NT Mutations in Splicing Figure 2. Analyzed sequences. The sequences are ordered according to the length of their longest PPS. Exons whose splicing was shown to depend on intact E+ 1 position are underlined. The PPS are singly underlined and other polypyrimidine stretches are dashed underlined. Sites of mutations are showed in bold. Seq. = sequence. (A) Sequences of the test set. (B) Sequences of the borderline set. doi: 10.1371 /journal .pone.0089570.g002 found only twice in total but in these cases the sequence was no more depicted as exonic (data not shown). Another factor that impacts the definition of an exon is the presence of SREs. These elements have been shown to be more active in proximity to splice sites [20] and therefore their overlap with the E + position cannot be ruled out. fn keeping with this view, a SRE spanning the 3'ss position has been already detected in the SMN1 and SMJV2 genes [21]. In order to evaluate the possibility that some G + 1 mutations affect splicing through changes in SRE elements rather than affecting U2AF35 binding to the splice sites we analyzed SRE predictions provided by Sroogle engine (see section material and methods for details) [16]. Using the approach that preserves the definition of exonic and intronic sequences prior to analysis, we observed four positive predictions among splicing-affecting mutations while there was only one such case among the non-affecting variants. Furthermore, when we indicated the whole sequence as exonic we detected more predictions of SRE changes among splicing-affecting mutations (eleven vs. two among the non-affecting variants), reaching statistical significance at p<0.01 (Fisher's exact test; Table S5 in File SI and Table S6). In conclusion, it appears that the disruption of SRE elements may, in parallel to intronic AG-dependence constitute another mechanism underlying splicing affection upon E + mutation (see discussion for further details). Expectedly, some overlap was observed between the resulting value ranges for the splicing-affecting and non-affecting mutations even in the parameters showing statistically significant differences. Nevertheless, we tried to define cut-off values that could assort the results of our predictions into two groups according to the splicing affection. To establish such cut-off values, we firstly searched for an approximate cut-off value that could discriminate these two types of sequences in the test set with minimal overlap. For integral parameters we employed exactly that value as a discriminative cutoff. For other parameters, we found the two closest non overlapping values from both splicing-affecting and non-affecting group and counted the final cut-off value as their arithmetic mean. The predicted scores and percentiles below these cut-off values and the differences between the predicted values for wild type and mutant sequences above the cut-off values are supposed to pertain to variants prone to affect splicing. Table 3 shows the selected cutoff values for the following variables: longest PPS, the number of pyrimidines in 25 nt upstream of the acceptor site, and the seven best discriminating in silico predicted parameters. In addition, for these same sequence parameters, Figure S2 shows the discriminative power of various potential cut-off values to discern between splicing affecting and non-affecting mutations in the test set sequences. A presumably splicing-affecting sequences presumably splicing non-affecting sequences BTK e10 BTKeW FASeS SeroGI e4 ATMeW ATM e32 CD55 e9 IL17RAe'\2 WßWe13 M (PID1) (PID2) (PID4) (PID3) (PID7) (PID6) (PID8) (PID9) (PID5) wt mut wt mut wt mut wt mut wt mut wt mut wt mut wt mut wt mut B ATP7A e16 BRCA2e12 BRCA2 e24 COL1A1 e11 M NF1 e24 (BOR5) (BOR2) (BOR3) (BOR1) (BOR4) wt mut wt mut wt mut wt mut wt mut Figure 3. Results of the splicing minigene analyses. RT-PCR analysis of the literature-derived E+ 1 variations. The splicing affecting sequences are underlined. (A) The test set sequences. cDNA bands originating from BTK exon 10 mutated minigene are numbered as follows: 1) cryptic 3'ss utilization 31 nt upstream of the authentic splice site (the aberrant exon starts at C.840-31G), 2) normally spliced RNA. (B) The borderline set sequences. doi:10.1371/journal.pone.0089570.g003 PLOS ONE I www.plosone.org 6 February 2014 | Volume 9 | Issue 2 | e89570 Exon 1st NT Mutations in Splicing 100 -i MaxEnt score diff. • • - ö S - 100MaxEnt pere. diff. • • PSSM score diff. oo 100-, NNSplice score diff.' • • -mo- il -oiB. —?o- 1.0-, 0.9- 0.8 0.7 0.6 PPT score (Kol) 0? -OH CO d d d d d p o d o • aj u- c # LT) J . J . J . o o o O) lO int m ö LA d LA d LA d LA d d •o d LA d d atedusin thediffe ity( calc sar u lO ^ OJ "C fN 'u 5 II -P •A -g •A -P g -P jp -.g a z ro m ro m O rN ro rN iitivityandspec elowthecut-off LO ifidence vO 00 CO 00 CO en 00 o\ iitivityandspec elowthecut-offc O O rv rs rv rs o O 3- c -Q 0 O o\ o\ o\ o p OJ vi u > d d d d d d vi CD í# fN 4 4 4 rN fN vO OJ \p -C c lO vD ro LA LA LA LA vO Jr QJ y 1" 0* d d d d d d d d d Jr QJ y 1"t/i y "~ al m * o 4-» .Q U C lO ra 1 / 1 tu fN i- "O T3 C II čP čP čP c QJ f— rv ó d 00 d d Oj d d 00 d vO d rs d l i s di 4 LA o LA LA LA rv 00 4 + E g lO ro rN m o v + i SS 0* — d d d d d d d d d C M_ „ on Ú o 134-1 C set rian OJ Ů 5 *^ £P T3 ^ VI > >, CU= O iepe ecific =30) sp sp sp sp íatura igina1 rtaint a z o o o o o o o o o ; o S ô lo rs LT) rs rv rs •o ro tr o cu a < 4- £ O (UP*-OJ 4-1 dl u 1 / 1 O) XI tS C D na c dl thete iuppo; 'E fid thete iuppo; c rv PM o o 00 •o O •O CT — „, 0 CO 00 rv p p en 00 C D d JIA > d d d d d d d 'vi LJ (CT3 í# di J . vo CO m o J . m D ra vi TJ U OJ4—» lO rN rv vo •o LA oJ >. 3 fT3 d d d d d d d d d C .ti (13 -C •rä £ > 4-» JD 'utA O OJ P tool ja tedvalues> vityandsp vethecut- to tivi Ô tedvalues> vityandsp vethecut- :— *w m tedvalues> vityandsp vethecut- c M U ' ŕ P ,c ai z LA o o LA LA o LA 00 red! ensi iab ,c m •o rs LT) o\ o\ CO •O red! ensi iab tu Q. vi Qj -C 4—* the The jene or 2 vi 5- M- OJ =S off3 if! f O # «0 # # m cording tedPPT utantse 3 val m 00 o d PS m m d cording tedPPT utantse 3 val 3 fN 00 00 o m o d PS m d proposedao ificiallymuta typeandm 0089570.t00 it= U *~ *~ «0 oó d fN *t fŇ m proposedao ificiallymuta typeandm 0089570.t00 cut-o ai u ence proposedao ificiallymuta typeandm 0089570.t00 SOQ c LA eren ence c.difference aiu 3J U 'ere| gart wild ione. Propo PPS Pyin2 T3 OJ O tscore tpere. rediffer c.difference differer differen valuesvv Dntainin luesfor ournal.c ro est «*•- 0 dl mu mu SCO per ore pere. i*_ uraT; p.Gly2281Val) in the exon 12 of BRCA2 gene and PID7 (c.1236 G>T; p.Trp412Cys) in the exon 10 o f ^ T M a r e recorded in the dbSNP database as polymorphisms with unknown frequency and clinical significance. The quality of the PPT depends on its overall length, its distance from the 3'ss, the number of pyrimidines and the number of consecutive uridines [22,23]. Notably, for the series of human 3'ss investigated by Fu et al. [3], the authors suggest that the length of the PPS represents the best discriminating feature between AGdependent and AG-independent introns; namely, the PPS was shown to encompass between 4 and 10 nt in the AG-dependent introns and between 9 and 16 nt in the AG-independent introns in the sequences analyzed in their experiments. This implies that the sequences with PPS of 10 nt or shorter have higher propensity to be aberrantly spliced upon E + mutation and deserve more detailed in vitro analysis [3]. When we reanalyzed these results after adding these 16 newly analyzed sequences to the Fu's original set, the PPS of AGdependent introns measured between 4 and 11 nt (with one outlier 18 nt long discussed below), and the PPS of AG-independent introns measured between 6 and 17 nt. Despite the significant difference between median values (8 nt for AG-dependent and 14 nt for AG-independent introns), the PPS length showed marked overlap. Notably, the AG-independent exons with shorter PPS generally had more than one polypyrimidine stretch in the PPT, indicating that even a few shorter stretches can assure AG- independence. Among the number of features that influence the exon definition, we would like to emphasize the importance of SREs. In this work, we also considered the predicted ESE disruption or PLOS ONE I www.plosone.org 9 February 2014 | Volume 9 | Issue 2 | e89570 Exon 1st NT Mutations in Splicing 2 £ "5 ESS creation spanning the E + position, as the events potentially E £ I affecting splicing. 5 2j From a mechanistic point of view, the ESE disruption and ESS .£ §" I creation differ significantly from each other. A protein factor _ .g £ bound to a newly created ESS (e.g. a protein from hnRNP family) S g _ could form a steric hindrance for U2AF35 binding [24]. Or, 3 D * possibly, such a protein could multimerize along the R N A and " in ^ block the binding of other factors, e.g. U2AF65, SF1, etc. [25]. In S __ _ fact, both these events were proposed in the splicing regulation of ;_ .£ _ SMM1 and SMM2 exon 7, where an ESS presence across the 3'ss g _ _ was shown [21]. Alternatively, ESS-bound protein could hamper the function of a downstream ESE [24]. All these events could "U £ JS >, Q. i S E ra .2 £ _TO"D "5 £ cu clearly impair the process of 3'ss recognition and cause the splicing aberration. However, the effects of disruption of an ESE spanning the 3'ss could be less clear, since the presence of ESE in this position may not seem logical at first glance. The reason is that most of the ESEbinding proteins were shown to facilitate 3'ss recognition through interaction with U2AF35 which can then bind to the 3'ss more easily (see [26] for review). However, it was shown that some ESEbinding SR-proteins can interact directly with the U2AF65 or with U2 snRNP [27,28], and for this reason these factors could £ !£ ^ promote splicing even in the absence of U2AF35. ° _ S In this respect, p54 protein from an SR-family that interacts j> _ 5! directly with the U2AF65 was suggested to be able to functionally > s 1 replace U2AF35 in its bridging other SR-proteins with U2AF65 ° o ai [27]. Might be that this or similar proteins could act in place of _ | U2AF35 in the 3'ss recognition that is essential for AG-dependent g 2 st introns. In concordance with this, not all the AG-dependent 0 £ ~?i introns were shown to depend on U2AF35 binding [3,29]. c 2 Therefore, disruption of binding site for such a protein could 5 S 5 impair U2AF65 binding to the PPT, thus precluding correct splice OJ "g £ site recognition. _ 15 Si The other possible function of an ESE spanning the 3 ss •goo. position might be the stabilization of U2 binding to the branch ^ ,£ point site. It was demonstrated that the RS-domain of an SR_ ^ _ protein contacts the BS in the prespliceosome and that such a | S _ contact helps in promoting the prespliceosome assembly [28,30]. •g 3 "S In fact, such an event occurs later in the spliceosome assembly 5 8 _ than the original recognition of 3'ss by the U2AF complex and is _;_ o. probably enabled by the lower association of U2AF with RNA in Si g £ the later spliceosomal complexes [31]. Therefore, disruption of an « « y ESE could harm the splicing independently on U2AF35 binding. •£ n c I1 1 addition, at least in AG-independent introns, an ESE spanning ° 5 g. the 3'ss could antagonize the function of proximal ESS, analogically to the situation in the SMN1 exon 7 [21]. In conclusion, therefore, we must consider the possibility that every E + mutation can affect not only the definition of a splice site, but may also introduce relevant changes in SRE elements that could lead to a splicing aberration even in the AG-independent splice sites. Furthermore, many described SREs are purine-rich j elements [32] which might provide an explanation as to why the G _ >. | g to A changes in the E + 1 position tend to be more easily tolerated | •_ 2 5 by the splicing machinery than G to T changes in our analysis. For S _ _ gj these reasons, we propose calling the sequences in which E + 1 ^ c § mutation affected splicing as "E+1 -dependent" and the other S | jg g sequences as "E+1 -independent" rather than AG-dependent or 1 g | _§- independent, unless AG-dependence is demonstrated. ^ £ ~R £ In keeping with this conclusion, SRE disruption has already .E -S s §• been proven in the case of the herein analyzed GH1 exon 3 £ 3 of m jj> mutation (FuO 1, c. 172 G>T) [33]. In addition, mutation c.517G> ^ £ ^ 5 T in the BRCA2 exon 7 (Evall) lies in a region where a splicing ^ "S § •§ enhancer was previously detected [34]. In fact, disruption of a functional splicing enhancer and/or creation of a splicing silencer E c a 10 February 2014 | Volume 9 | Issue 2 | e89570 Exon 1st NT Mutations in Splicing could provide an explanation for the cases of splicing affection in the introns with a good quality PPS and strong splice sites, such as in the cases of ATM exon 10, BRCA2 exon 12 and CFTR exon 4 mutations (designated as PID7, BOR2 and EVAL4, respectively). Accordingly, potentially harmful SRE changes were predicted in all these three cases. Furthermore, we detected significantly higher occurrence of SRE changes between E + -dependent sequences than in the E+1 -independent ones. Nonetheless, the predictions of the SREs encompassing the E + 1 position were only partly helpful in discerning the E+1 -dependent from E+1 -independent sequences. In fact, although we detected a difference between the number of positive predictions of supposedly relevant SRE changes between the two groups of E + 1 dependent and E+1 -independent sequences, we suppose that these predictions overestimated the number of functional SRE changes. This is quite a common occurrence and it has been described several times that SRE predictions suffered from a high number of false positive and false negative results [35,36]. Most importantly, as splicing dependence on the intact E + 1 position is difficult to predict from sequence inspection alone, we examined the ability of available software tools to predict exonic E + 1 dependence. Interestingly, we obtained the most promising results using the tools that estimated the overall quality of the splice site (i.e., MaxEnt, PSSM and NNSplice). The best predictions were obtained with the MaxEnt program, which proved to be highly efficient in previous studies as well [37,38]. These results are consistent with the MaxEnt program basing its predictions on calculations of the interdependence of non-adjacent positions, which may correspond well to the cooperation of several signal sequences during splice site recognition [9]. In contrast, the outcomes of PPT predictions were not optimal. In fact, the predicted value ranges between E+1 -dependent and E+1 -independent groups of sequences overlapped considerably. The BS predictions produced similar results, which may correspond to the looser connection between the quality of the BS and the AG-dependence of the 3'ss [22]. However, the low performance of the PPT predictors is more remarkable because this element is believed to control the intron AG-dependence and, as explained above, our results confirm it to be major determinant of E + 1 dependence. The most likely explanation is that this tool does not take into account the cooperation of factors bound to PPT and to other elements such as the splice site, BS or SREs [26,39,40]. Finally, in order to facilitate the work of genetic diagnosticians who often struggle with unclassified sequence variants whose action on the splicing process may be unclear, we have used our results to define discriminating border values for the software tools that might help assess whether a variant will be likely to influence splicing or not. Other authors have proposed that a 10% difference between the wild type- and mutant-predicted scores is an acceptable cut-off value [41,42]. However, this value is neither suitable for nor applicable to all of the instruments tested (e.g., PSSM, see Tables S2, S3 and S4 in File SI). Although other authors have proposed that cut-off values might be calculated to become gene specific [41], an alternative could be to obtain cut-off values that might be specific for a particular splicing signal and the position of the mutation. In accordance to this view, for our sets of G + 1 mutations, we have proposed specific cut-off values that will allow E+1 -dependent from E+1 -independent sequences to be distinguished with reasonably high sensitivity and specificity. An exception was the case of BOR4 mutation {Ml exon 24; see Table S2 in File SI) that was predicted as splicing-affecting by most of the tools, while it did not affect splicing in fact. A possible explanation for this failure could be a rather specific, 14-nt long PPS located quite far from the 3'ss that may lead the prediction tools to mistakenly consider this splice site as weak. Nonetheless, owing to different underlying algorithms, individual prediction tools have distinct advantages and disadvantages. Therefore, from our analysis we have further extracted three combinations of prediction parameters that when taken together, perform even better than the individual cut-off values. Furthermore, in addition to their good results when evaluated with the set of sequences with artificially mutated PPT, these combined predictions performed well even with the natural sequence variants. For example, in the set of 9 evaluation sequences, the length of the PPS itself would have lead to misprediction of E+1 -dependence in 7 cases, and even the individual in silico tools predictions would have failed in 19% of the predictions (Table S7 in File SI). However, when we applied combined predictions the correct splicing outcome could be predicted in all cases. We therefore suggest that the proposed individual MaxEnt and PSSM cut-off values and their combinations should be considered when evaluating the potential splicing effects of new G + 1 mutations. To facilitate work using the cut-off limits, we now add a simple calculation table as a supplement to this article (Table S8). In conclusion, therefore, we believe that this work could serve as a preliminary guide for estimating the impact of exon first nucleotide mutations on splicing. Supporting Information Figure SI Complete list of the sequences used for evaluation of in silico prediction tools effectivity. Exons which splicing was shown to depend on intact E + 1 position are underlined. The PPS are simply underlined, other polypyrimidine stretches are dashed underlined. Sites of mutations are showed in bold. Note that this figure contains all the herein used sequences except from artificially mutated ones of the Fu-mut sets, i.e. the test set, borderline set and the evaluation set. (TIFF) Figure S2 Discriminative power of various potential cut-off values. The charts show percentage of correctly sorted test set sequences according to mutation induced splicing affection (success rate) plotted against various cut-off values of individual sequence parameters. The dashed line and the number beside show finally selected cut-off value (selected according to the rules explained in results part). Note that the predicted scores and percentiles below these cut-off values and the differences between the predicted values for wild type and mutant sequences above the cut-off values are supposed to pertain to variants prone to affect splicing. Diff. = difference, perc. = percentile, Py25 = number of pyrimidines in the 25 nucleotides upstream from splice site. (TIF) Figure S3 Sequences of the evaluation set. Exons which splicing was shown to depend on intact E + 1 position are underlined. The PPS are simply underlined, other polypyrimidine stretches are dashed underlined. Sites of mutations are showed in bold. (TIFF) Table S6 Individual SRE predictions. The individual predictions for all the herein analyzed sequences obtained at Sroogle engine. The table is divided into two parts, one with borders of exons and introns being kept when inserted in the Sroogle engine; the other with the whole sequence considered as an exonic one. Mut = mutant, wt = wild type. Other abbreviations PLOS ONE I www.plosone.org 11 February 2014 | Volume 9 | Issue 2 | e89570 Exon 1st NT Mutations in Splicing as well as the colors identifying SRE types (red for enhancers, green for silencers and gray for nonspecified regulators) were adopted from Sroogle [S4]. (XLSX) Table S8 Ready-set formulae to facilitate making conclusions of the individual and the combinatorial predictions. Acc = acceptor splice site; Longest Py stretch = the longest uninterrupted polypyrimidine stretch; Py = number of pyrimidines (in 25 upstream from 3'ss). (XLS) File SI Includes Tables S1-S5 and S7. Table SI. List of analyzed E + 1 mutations of the test set and the borderline set. Exon skipping analyses were described in this work for mutations 013, HK08, all the PID and all the B O R sequences, the rest was adopted from Fu et al. [SI]. The sequences that presented aberrant splicing following mutation are highlighted in light orange. The exons highlighted in yellow were selected as presumably splicing-affecting (the length of their PPS being 10 nt at maximum, containing the uninterrupted T-stretch not longer than 3 nt) and those highlighted in green as presumably splicing non-affecting (the minimal length of their PPS being 11 nt, with Tstretches of 4 nt at minimum; see results section). The depicted protein change is a predicted change based solely on DNA-level knowledge. The number of pyrimidines upstream from the 3'ss was counted in 25- (50-) nt sequences. The values that do not fall (according to the mutation influence on splicing) into the range of herein proposed cut-off limits are marked in blue. BOR = sequences from the "borderline set" of mutations. Py = number of pyrimidines, PPS = the longest uninterrupted polypyrimidine stretch, RefSeq = reference sequence in the NCBI database, skip. = ratio of exon skipping in the minigene analyses of wild type (wt) or mutant (mut) sequences. * there are several discrepancies in the numbering of exons in the reference sequences compared to the numbering shown in Fu et al. (2011): ETA1 exon 10 is depicted as number 12 in the reference sequence; CAPM3 exons 10 and 17 are depicted as numbers 5 and 12 in the reference sequence, respectively. ** mutation C.1350OT has not been reported yet. *** the mutation was derived from RAPID database where mutations of G to A, C and T were depicted in the same position. However, the change selected for this article, G>A, was later found to be mistakenly obtained from [S2], where a mutation at adjacent position was described. Table S2. Predicted values for the E + 1 mutations using instruments evaluating the overall strength of the 3'splice site. The values that do not fall (according to their sequences influence on splicing) into the range of the herein proposed cut-off limits are marked in blue. The sequences that were shown to adopt aberrant splicing upon mutation are highlighted in light orange. Diff. = difference, perc. = percentile, seq. = sequence Table S3. Predicted values for the PPT of the E + 1 mutated sequences. The sequences that were shown to adopt aberrant splicing upon mutation are highlighted in light orange. "—" indicates the cases where the computer tool gave no values. Perc. = percentile, dist. = distance Table S4. Predicted values for the BS of the E + 1 mutated sequences. The sequences that were shown to adopt aberrant splicing upon mutation are highlighted in light orange. "—" indicates the cases where the computer tool gave no values. Perc. = percentile, dist. = distance Table S5. Prediction of SRE changes (using Sroogle engine). The sequences that were shown to adopt aberrant splicing upon mutation are highlighted in light orange. Positive predictions are marked in green. For an explanation, see material and methods. The statistical comparison of the SRE changes in splicing-alfecting and non-affecting samples was counted only from the test-set of sequences (i.e. without the borderline set sequences). Table S7. Combined predictions of splicing affection on nine evaluation sequences. a Each combined prediction was considered as positive if two (or more) of the three predicted values exceeded the herein proposed cut-off values of the individual tools. The individual values that do not fall into the range of herein proposed cut-off limits are marked in blue. Predictions being in accordance with detected splicing affection are marked in green, the discrepancies are in orange. Py25 = number of pyrimidines in the 25 nucleotides upstream from splice site; M E s.d. — difference between wild type and mutant sequence scores predicted by MaxEnt program; M E p.d. = difference between wild type and mutant sequence percentiles predicted by MaxEnt program; PSSM s.d.: accordingly. (DOC) Acknowledgments We are grateful to Prof. Litzman, Dr. Rihak, Dr. Kralickova and Dr. Skopkova for providing the patient samples. We thank L. Kazdova, M . Plotena and C. Stuani for their technical assistance. Author Contributions Conceived and designed the experiments: L G PL BR EB FEB TF. Performed the experiments: L G PL BR. Analyzed the data: L G EB L D TF. Contributed reagents/materials/analysis tools: EB FEB LD TF. Wrote the paper: L G PL BR EB FEB L D TF. References 1. Hastings M L , Kramer A R (2001) Pre-mRNA splicing in the new millennium, Curr Opin Cell Biol 13: 302 309. 2. Baralle D, Baralle M (2005) Splicing in action: assessing disease causing sequence changes. J Med Genet 42: 737-748, 3. Fu Y, Masuda A, Ito M , ShinmiJ, Ohno K (2011) AG-dependent 3'-splice sites are predisposed to aberrant splicing due to a mutation at the first nucleotide of an exon. Nucleic Acids Res 39: 4396-4404, 4. Reed R (1989) The organization of 3' splice-site sequences in mammalian introns. Genes Dev 3: 2113 2123. 5. Guth S, Martinez C, Gaur R K , Valcarcel J (1999) Evidence for substratespecific requirement of the splicing factor U2AF(35) and for its function after polypyrimidine tract recognition by U2AF(65). Mol Cell Biol 19: 8263 8271, 6. Wu S, Romfo C M , Nilsen T W , Green M R (1999) Functional recognition of the 3' splice site A G by the splicing factor U2AF35. Nature 402: 832 835. 7. Shapiro M B , Senapathy P (1987) R N A splice junctions of different classes of eukaryotes: sequence statistics and functional implications in gene expression. Nucleic Acids Res 15: 7155 7174. 8. Reese M G , Eeckman F H , Kulp D, Haussler D (1997) Improved splice site detection in Genie. J Comput Biol 4: 311-323, 9. Yeo G , Bürge C B (2004) Maximum entropy modeling of short sequence motifs with applications to R N A splicing signals. J Comput Biol 11: 377 394, 10. Kol G, Lev-Maor G , Ast G (2005) Human-mouse comparative analysis reveals that branch-site plasticity contributes to splicing regulation. Hum Mol Genet 14: 1559 1568. 11. Schwartz SH, Silva J, Burstein D, Pupko T, Eyras E, et al. (2008) Large-scale comparative analysis of splicing signals and their corresponding splicing factors in eukaryotes. Genome Res 18: 88 103, 12. Fairbrother W G , Yeh RF, Sharp PA, Bürge CB (2002) Predictive identification of exonic splicing enhancers in human genes. Science 297: 1007-1013, 13. Zhang X H , Chasin L A (2004) Computational definition of sequence motifs governing constitutive exon splicing. Genes Dev 18: 1241 1250, 14. Cartegni L, WangJ, Zhu Z, Zhang M Q , Krainer A R (2003) ESEfinder: A web resource to identify exonic splicing enhancers. Nucleic Acids Res 31: 3568 35 71, 15. Wang Z, Rolish M E , Yeo G , Tung V , Mawson M , et al. (2004) Systematic identification and analysis of exonic splicing silencers. Cell 119: 831 845, 16. Schwartz S, Hall E, Ast G (2009) S R O O G L E : webserver for integrative, userfriendly visualization of splicing signals. Nucleic Acids Res 37: W189 192, PLOS ONE I www.plosone.org 12 February 2014 | Volume 9 | Issue 2 | e89570 Exon 1st NT Mutations in Splicing 17. Sickmier EA, Frato K E , Shen H , Paranawithana SR, Green M R , et al. (2006) Structural basis for polypyrimidine tract recognition by the essential pre-mRNA splicing factor U2AF65. Mol Cell 23: 49 59. 18. Mackereth C D , Madl T, Bonnal S, Simon B, Zanier K , et al. (2011) Multidomain conformational selection underlies pre-mRNA splicing regulation by U2AF. Nature 475: 408 411. 19. Jenkins JL, Agrawal AA, Gupta A, Green M R , Kielkopf C L (2013) U2AF65 adapts to diverse pre-mRNA splice sites through conformational selection of specific and promiscuous RNA recognition motifs. Nucleic Acids Res 41: 3859 3873. 20. Graveley BR, Hertel KJ, Maniatis T (1998) A systematic analysis of the factors that determine the strength of pre-mRNA splicing enhancers. E M B O J 17: 6747 6756. 21. Doktor T K , Schroeder LD, Vested A, Palmfeldt J , Andersen HS, et al. (2011) SMN2 exon 7 splicing is inhibited by binding of hnRNP A l to a common ESS motif that spans the 3' splice site. Hum Mutat 32: 220-230. 22. Roscigno RF, Weiner M , Garcia-Blanco M A (1993) A mutational analysis of the polypyrimidine tract of introns. Effects of sequence differences in pyrimidine tracts on splicing. J Biol Chem 268: 11222 11229. 23. Coolidge CJ, Seely RJ, Patton J G (1997) Functional analysis of the polypyrimidine tract in pre-mRNA splicing. Nucleic Acids Res 25: 888 896. 24. Chen M , Manley J L (2009) Mechanisms of alternative splicing regulation: insights from molecular and genomics approaches. Nat Rev Mol Cell Biol 10: 741 754. 25. Spellman R, Smith C W (2006) Novel modes of splicing repression by PTB, Trends Biochem Sci 31: 73-76, 26. Graveley BR, Hertel KJ, Maniatis T (2001) The role of U2AF35 and U2AF65 in enhancer-dependent splicing. R N A 7: 806 818, 27. Zhang WJ, W u J Y (1996) Functional properties of p54, a novel SR protein active in constitutive and alternative splicing. Mol Cell Biol 16: 5400 5408, 28. Shen H , Green M R (2004) A pathway of sequential arginine-serine-rich domainsplicing signal interactions during mammalian spliceosome assembly. Mol Cell 16: 363-373. 29. Pacheco TR, Coelho M B , Desterro J M , Mollet I, Carmo-Fonseca M (2006) In vivo requirement of the small subunit of U2AF for recognition of a weak 3' splice site. M o l Cell Biol 26: 8183 8190. 30. Shen H , Kan J L , Green M R (2004) Arginine-serine-rich domains bound at splicing enhancers contact the branchpoint to promote prespliceosome assembly, Mol Cell 13: 367 376. 31. Bennett M , Michaud S, Kingston J , Reed R (1992) Protein components specifically associated with prespliceosome and spliceosome complexes. Genes Dev 6: 1986 2000. 32. Cartegni L, Chew SL, Krainer A R (2002) Listening to silence and understanding nonsense: exonic mutations that affect splicing. Nat Rev Genet 3: 285 298, 33. Shariat N , Holladay C D , Cleary R K , Phillips JA, Patton J G (2008) Isolated growth hormone deficiency type II caused by a point mutation that alters both splice site strength and splicing enhancer function. Clin Genet 74: 539 545, 34. Gaildrat P, Krieger S, D i Giacomo D, AbdatJ, Révillion F, et al. (2012) Multiple sequence variants of BRCA2 exon 7 alter splicing regulation. J Med Genet 49: 609 617. 35. Auclair J , Busine M P , Navarro C, Ruano E, Montmain G, et al. (2006) Systematic mRNA analysis for the effect of M L H 1 and MSH2 missense and silent mutations on aberrant splicing. Hum Mutat 27: 145 154, 36. Lastella P, Surdo N C , Resta N , Guanti G, Stella A (2006) In silico and in vivo splicing analysis of MLH1 and MSH2 missense mutations shows exon- and tissue-specific effects. B M C Genomics 7: 243, 37. Vorechovský I (2006) Aberrant 3' splice sites in human disease genes: mutation pattern, nucleotide structure and comparison of computational tools that predict their utilization. Nucleic Acids Res 34: 4630-4641. 38. Buratti E, Chivers M , Královicova J , Romano M , Baralle M , et al. (2007) Aberrant 5' splice sites in human disease genes: mutation pattern, nucleotide structure and comparison of computational tools that predict their utilization, Nucleic Acids Res 35: 4250 4263. 39. Abovich N , Rosbash M (1997) Cross-intron bridging interactions in the yeast commitment complex are conserved in mammals. Cell 89: 403 412, 40. Berglund JA, Abovich N , Rosbash M (1998) A cooperative interaction between U2AF65 and mBBP/SFl facilitates branchpoint region recognition. Genes Dev 12: 858 867. 41. Houdayer C, Dehainault C, Mattler C, Michaux D, Caux-Moncoutier V , et al, (2008) Evaluation of in silico splice tools for decision-making in molecular diagnosis. Hum Mutat 29: 975 982. 42. Zampieri S, Buratti E, Dominissini S, Montalvo A L , Pittis M G , et al. (2011) Splicing mutations in glycogen-storage disease type II: evaluation of the full spectrum of mutations and their relation to patients' phenotypes. Eur J Hum Genet 19: 422 431. PLOS ONE I www.plosone.org 13 February 2014 | Volume 9 | Issue 2 | e89570 3.2. Splicing-affecting variants in selected PID-related genes Grodecká L, Hujová P, Kramářek M , et al. Systematic analysis of splicing defects in selected primary immunodeficiencies-related genes. Submitted. Huge variance exists in the assessment of how prevalent are the splicing-affecting mutations in general. While some findings propose up to 50 % fraction of all disease-causing mutations, others are much more conservative, coming as low as 15 % (which usually covers only splice site affecting mutations, not those in SREs).3 0 , 1 2 0 , 1 7 8 , 2 2 1 This situation reflects a highly variable prevalence of splicing affecting mutations in different genes, which is apparently influenced by two facts: i) splicing affecting mutations seem to be especially abundant in genes containing high number of exons1 0 8 , 2 2 2 and ii) alternative exons tend to be more susceptible to splicing-affecting variants.97 Regarding the PID-related genes, many individual findings on splicing variants were described, but no study evaluated these variants in general. Yet, with the expansion of the high throughput sequencing, the knowledge of splicing affection, its prevalence and the locations of essential splicing regulatory elements becomes crucial for proper assessment of variants' pathogenic potential. In order to tentatively reduce the portion of VUS in the PID-related genes, we have systematically examined 92 variants found in 6 PID-related genes (BTK, CD40LG, IL2RG, SERP1NGJ, STATS, and WAS) for splicing defects. The results of these analyses showed that: Twenty out of 92 in silico I 55 in vitro tested variants (21.7 %) resulted in aberrant splicing. Fourteen of these samples disrupted and 2 created novel splice sites, the remaining 4 most probably affected SREs. These results conform to the more conservative estimates of the splicing affection prevalence, possibly due to the fact that the number of reliably established alternatively spliced exons in the herein inspected genes is rather limited. In addition, while the splice site affections mostly led to massive splicing aberration, the changes of SREs shifted the proportion of alternative splicing by 5-25 %. Therefore, their relevance for disease development remains uncertain. Thirty out of 32 RT-PCR results from patients' blood-derived R N A confirmed the existence or absence of splicing affection originally detected in the minigene settings. The two discordant samples were positive for splicing aberration in the minigene 62 while negative in blood. We suppose that the discord resulted from our inability to detect the aberrantly spliced transcripts in the patients' blood, in one case due to possible NMD-based degradation, in the other case due to overall low amount of the specific mRNA. Overall, we showed reasonably good reliability of the minigene assay, at least in terms of the absence/existence of splicing affection. Admittedly, its concordance with the blood-derived R N A was lower when looking at particular aberrant splicing patterns (8 samples out of 10 positive for splicing affection in blood showed different splicing pattern in blood and minigene). Control of endogenous splicing of the inspected genes in the used cell lines may demonstrate unexpected results and might pinpoint gene regions that could be tested in the minigenes with higher/lower reliability. As the alternative splicing patterns as well as particular outcome of splicing aberration are often tissue specific,223 we have tentatively explored the endogenous expression of the analysed genes in the cell-lines used for our minigene assay. The overall expression of these genes mostly corresponded to the data derived from literature and databases. On the other hand, we have detected several differences, some of which indicated the changed splicing pattern of the gene (data not shown). Overall, such data might form an interesting complement to the minigene assay (which is still rather artificial method) and to indicate which of the tested exons/minigenes show higher or less reliable results. An interesting and possibly useful pattern could be found among predicted cryptic splice sites. So far, forecast of splicing outcome after splice site disruption is a difficult task for most prediction tools (see chapter 1.1.4.1 'In silico predictions'). Using general splice site predictors (MaxEnt and NNSplice), we showed that the activated cryptic splice sites in exons tended to be the closest to their authentic counterparts, while the intronic activated splice sites were the strongest among the predicted. We believe that this pattern might mirror the general ESE/ESS ratio, in introns being opposite than in exons. Hopefully these results may be further proved on a larger data set and eventually could contribute to the development of more accurate prediction models on cryptic splice site activation. Many SRE-predicting tools were able to globally detect SRE changes, while only several newer predictors provided potentially useful forecasts on splicing affection for the particular variants. In more detail, ESR-seq score, AHZE i score and EX-SKIP showed score changes that corresponded to the observed splicing aberrations in 4, 4 63 and 3 cases, respectively. On the other hand, the specificity of these tools is still very limited. It follows that these tools could still serve just as a crude filter for the preliminary selections of variants to be further tested in vitro, until more reliable, possibly more complex predictions taking into account most of the influencing factors. 64 3.3. Splice site competition leads to introduction of multiple pseudoexons in the IDS gene Grodecká L, Kováčova T, Kramářek M, et al. Detailed molecular characterization of a novel IDS exonic mutation associated with multiple pseudoexon activation. J Mol Med Berl Ger. November 2016. doi: 10.1007/sOO 109-016-1484-2. As already indicated, one of the huge complications for the genetic diagnostics when a variant is likely to affect splicing is the difficulty of predicting the aberrant spicing outcome, which can be sometimes truly surprising. This absolutely applies to the c.l092T>A variant found in the IDS gene of a patient suffering from mild form of mucopolysaccharidosis type II. While being silent at the protein level, this exonic variant introduced moderately strong de novo splice site, so the splicing affection was highly probable. Strikingly, RT-PCR from patient's fibroblast showed presence of multiple aberrant transcripts, most of which incorporated various internal parts of upstream intron. In other words, the c.l092T>A led to the activation of multiple pseudoexons. At the beginning of this project, the known mechanisms resulting in pseudoexon activation included: i) creation of a new splice site;1 1 0 1 1 1 ii) change of a SRE inside or near to the pseudoexonic sequence;29,87 1 1 3 iii) disruption of an authentic splice site in the neighbouring exon;1 1 4 1 1 5 1 1 7 and iv) probable ESE disruption in the neighbouring exon, leading to the bypass of authentic splice site and to pseudoexon activation.118 With regard to these options, our first hypothesis that the variant c.l092T>A might have disrupted or created an SRE, together with the novel splice site introduction. However, multiple lines of evidence indicated that: The hypothesised SRE change is not the major cause of the observed splicing pattern, despite we can not totally exclude its minor effect The main cause of the splicing defect is the competition between de novo and the authentic splice site. Several experiments using minigenes with either weakened de novo splice site or strengthened authentic splice site showed gradual correction of the splicing pattern. In addition, sole disruption of the authentic splice site (together with its neighbouring cryptic splice site) resulted in multiple pseudoexon activation. The proposed mechanism of splicing aberration observed in the case of c.l092T>A mutation is that the de novo splice site outcompetes the authentic splice site for the 86 spliceosomal recognition and that this bypass then allows the pseudoexonic splice sites to be effectively recognized and selected for splicing. Of note, the activated pseudoexons overlap with the sequence of IDS alternative terminal exon, which acceptor splice site was used by several of the activated pseudoexons. We believe that the existence of alternative terminal exon could facilitate the pseudoexon activation. We observed a previously unreported mechanism of multiple pseudoexon inclusion. Although it might be very rare, it is important to take this mechanism into account when dealing with mutations leading to splice site bypass, especially when an alternative terminal exon is present in the neighbouring sequence. 87 3.4. Variant c.4211-32_-13del in the FBN1 gene leads to major exon skipping Šípek A, Grodecká L, Baxová A, et al. Novel FBN1 gene mutation and maternal germinal mosaicism as the cause of neonatal form of Marfan syndrome. Am J Med Genet A. 2014; 164A(6): 1559-1564. doi: 10.1002/ajmg.a.36480. Here we have presented a clinical report on a girl with severe neonatal form of Marfan syndrome that was highly suspicious for a splicing defect. A novel variant c.4211-32_-13del was found in the patient's FBN1 gene. This variant markedly impaired the acceptor splice site of exon 35, judging by in silico prediction results. Using the minigene assay, we showed that the mutation results in major exon 35 skipping, confirming the presumed splicing affection. However, vast majority of mutations leading to the infancy-onset Marfan syndrome are located in exons 24-32.2 2 4 - 2 2 6 Therefore, we speculated that this mutation might either induce multiple exon skipping including several upstream exons or it may be one of the rare cases of severe-form-causing mutations outside the region of exons 24-32. In addition, the same mutation was found in the patient's mother, but only in a minor portion of her blood cells, indicating probable germline mosaicism. Despite the mother was originally considered as asymptomatic, more detailed examination identified minor echocardiographic findings and a unilateral lens ectopy, phenotypic features connected to Marfan syndrome. Therefore, the knowledge of splicing affection by the c.4211-32_-13del variant is highly important also for other pregnancies of the mother, including potential prenatal or preimplantation diagnostics. 99 3.5. Variant C.2440-6OG in the CDH1 gene does not markedly affect splicing Grodecká L, Kramářek M, Lockerová P, et al. No major effect of the CDH1 C.2440-6OG mutation on splicing detected in last exon-specific splicing minigene assay. Genes Chromosomes Cancer. 2014;53(9):798-801. doi:10.1002/gcc.22186. Here, we have investigated C.2440-6OG CDH1 variant that was suspected for predisposing its carriers to the development of diffuse gastric cancer. In fact, there were discrepant data concerning this subject: More et al.2 2 7 described variant carrier with aberrant splicing detected using RT-PCR, while Molinaro et al.2 2 8 found no effect on splicing in another carrier of the same variant. We have investigated another carrier who showed no splicing affection using RT-PCR from blood-derived RNA. When we employed our minigene system, we detected only negligible splicing aberration which was in fact most probably minigene-specific. In addition, we showed that it was rather improbable that this variant has the pathogenic potential due to its high prevalence in general population. We hypothesised that its previously reported effect on splicing might have been caused either by a specific haplotype or genomic milieu of the inspected carriers. Of note, we have adjusted our minigene system for this assay to make it able to reliably reflect the splicing changes of the terminal exons. 106 4. Conclusions The initial goal of my research was to find out the prevalence of splicing-affecting variants among the PID-related genes and to optimize the way to recognize them. For this purpose, 111 individual gene variants (mostly found in the PID-related genes) have been analysed in silico, and 64 of them in vitro using splicing minigene assay and RT-PCR of the patients'-derived RNA, when available. The general prevalence of splicing-affecting variants in the PID-related genes was found 21.7 %, with only a minor representation of SRE-affecting variants. A recommendation follows that in PID-diagnostic settings, a variant should be inspected for a SRE affection only when multiple a priori evidence indicates that splicing affection through a SRE change is probable and may lead to the observed phenotype or the expected disorder that might manifest in future. Comparison of the two in vitro methods showed high reliability of the minigene assay in terms of presence or absence of a splicing affection. On the other hand, the concordance was lower when looking at the particular aberrant splicing patterns, making transmission of such results into the diagnostic recommendations somewhat difficult. Still, this assay may be very useful especially when patient's R N A is not available and other lines of evidence highly support the pathogenic potential of an inspected variant. Hence, I support the endeavour to embody this method more firmly into the diagnostic settings.122 Evaluation of splicing in silico predictors enabled us to tentatively define cut-off values helping discriminate splicing-affecting from non-affecting variants located at the first nucleotide position of exons. Further, the data showed promising results of cryptic splice site predictions and of SREs, both of which still turn out difficult for the prediction tools. All these outcomes may delineate future path both for prediction tools developers as well as for their general users. In addition, the results showed a novel mechanism of multiple pseudoexon activation and indicated possible new mechanism of 3'ss recognition. Overall, I believe that these outcomes extend our knowledge about splicing aberration and its detection. Some of them directly influenced diagnostic decisions, some may have more or less direct impact on the genetic diagnostics in future. I l l 5. References 1. Fu X-D. Towards a Splicing Code. Cell. 2004;119(6):736-738. doi:10.1016/j.cell.2004.11.039. 2. Matlin AJ, Clark F, Smith CWJ. Understanding alternative splicing: towards a cellular code. Nat Rev Mol Cell Biol. 2005;6(5):386-398. doi:10.1038/nrml645. 3. Brown TA. Genomes. 2nd ed. Oxford: Wiley-Liss; 2002. http://www.ncbi.nlm.nih.gov/books/NBK21128/. Accessed November 7, 2016. 4. Hube F, Francastel C. Mammalian Introns: When the Junk Generates Molecular Diversity. IntJMolSci. 2015;16(3):4429-4452. doi:10.3390/ijmsl6034429. 5. Derrien T, Johnson R, Bussotti G, et al. The GENCODE v7 catalog of human long noncoding RNAs: Analysis of their gene structure, evolution, and expression. Genome Res. 2012;22(9): 1775-1789. doi: 10.1101/gr. 132159.111. 6. The FANTOM Consortium. The Transcriptional Landscape of the Mammalian Genome. Science. 2005;309(5740):1559-1563. doi:10.1126/science.H12014. 7. Hastings ML, Krainer AR. Pre-mRNA splicing in the new millennium. Curr Opin Cell Biol. 2001;13(3):302-309. 8. Chen M , Manley JL. Mechanisms of alternative splicing regulation: insights from molecular and genomics approaches. Nat Rev Mol Cell Biol. September 2009. doi:10.1038/nrm2777. 9. Will CL, Luhrmann R. Spliceosome Structure and Function. Cold Spring Harb PerspectBiol. 2011;3(7):a003707-a003707. doi:10.1101/cshperspect.a003707. 10. De Conti L, Baralle M , Buratti E. Exon and intron definition in pre-mRNA splicing. Wiley Interdiscip Rev RNA. 2013;4(l):49-60. doi:10.1002/wrna.H40. 11. Schneider M, Will CL, Anokhina M , Tazi J, Urlaub H, Lührmann R. Exon Definition Complexes Contain the Tri-snRNP and Can Be Directly Converted into B-like Precatalytic Splicing Complexes. Mol Cell. 2010;38(2):223-235. doi:10.1016/j.molcel.2010.02.027. 12. Turunen JJ, Niemelä EH, Verma B, Frilander MJ. The significant other: splicing by the minor spliceosome. Wiley Interdiscip Rev RNA. 2013;4(l):61-76. doi:10.1002/wrna.H41. 13. Fu X-D, Ares M . Context-dependent control of alternative splicing by RNA-binding proteins. Nat Rev Genet. 2014;15(10):689-701. doi:10.1038/nrg3778. 14. Kalsotra A, Cooper TA. Functional consequences of developmentally regulated alternative splicing. Nat Rev Genet. 2011;12(10):715-729. doi:10.1038/nrg3052. 15. Melamed Z, Levy A, Ashwal-Fluss R, et al. Alternative Splicing Regulates Biogenesis of miRNAs Located across Exon-Intron Junctions. Mol Cell. 2013;50(6):869-881. 112 doi:10.1016/j.molcel.2013.05.007. 16. Lee Y, Rio DC. Mechanisms and Regulation of Alternative Pre-mRNA Splicing. Annu RevBiochem. 2015;84(l):291-323. doi:10.1146/annurev-biochem-060614-034316. 17. Faustino NA. Pre-mRNA splicing and human disease. Genes Dev. 2003;17(4):419-437. doi: 10.1101/gad. 1048803. 18. Biamonti G, Caceres JF. Cellular stress and R N A splicing. Trends Biochem Sei. 2009;34(3):146-153. doi: 10.1016/j.tibs.2008.11.004. 19. Holste D, Ohler U. Strategies for Identifying R N A Splicing Regulatory Motifs and Predicting Alternative Splicing Events. PLoS Comput Biol. 2008;4(l):e21. doi:10.1371/journal.pcbi.0040021. 20. Sohail M , Xie J. Diverse regulation of 3' splice site usage. Cell Mol Life Sei. 2015;72(24):4771-4793. doi:10.1007/s00018-015-2037-5. 21. Zhang MQ. Statistical features of human exons and their flanking regions. Hum Mol Genet. 1998;7(5):919-932. 22. Bürge CB, Padgett RA, Sharp PA. Evolutionary fates and origins of U12-type introns. Mol Cell. 1998;2(6):773-785. 23. Wu Q, Krainer AR. AT-AC pre-mRNA splicing mechanisms and conservation of minor introns in voltage-gated ion channel genes. Mol Cell Biol. 1999;19(5):3225-3236. 24. Szafranski K, Schindler S, Taudien S, et al. Violating the splicing rules: TG dinucleotides function as alternative 3' splice sites in U2-dependent introns. Genome Biol. 2007;8(8):R154. doi:10.1186/gb-2007-8-8-rl54. 25. Sun H, Chasin LA. Multiple splicing defects in an intronic false exon. Mol Cell Biol. 2000;20(17):6414-6425. 26. Cartegni L, Chew SL, Krainer AR. LISTENING TO SILENCE A N D UNDERSTANDING NONSENSE: EXONIC MUTATIONS THAT AFFECT SPLICING Nat Rev Genet. 2002;3(4):285-298. doi:10.1038/nrg775. 27. Padgett RA. New connections between splicing and human disease. Trends Genet. 2012;28(4): 147-154. doi: 10.1016/j tig.2012.01.001. 28. Baralle D. Splicing in action: assessing disease causing sequence changes. J Med Genet. 2005;42(10):737-748. doi:10.1136/jmg.2004.029538. 29. Pagani F, Stuani C, Tzetis M , et al. New type of disease causing mutations: the example of the composite exonic regulatory elements of splicing in CFTR exon 12. Hum Mol Genet. 2003; 12(10): 1111-1120. 30. Buratti E. Defective splicing, disease and therapy: searching for master checkpoints in exon definition. Nucleic Acids Res. 2006;34(12):3494-3510. doi:10.1093/nar/gkl498. 31. Martinez-Contreras R, Cloutier P, Shkreta L, Fisette J-F, Revil T, Chabot B. hnRNP 113 proteins and splicing control. Adv Exp MedBiol. 2007',623:123-147'. 32. Wang Z, Bürge CB. Splicing regulation: From a parts list of regulatory elements to an integrated splicing code. RNA. 2008;14(5):802-813. doi:10.1261/rna.876308. 33. Kanopka A, Mühlemann O, Akusjärvi G. Inhibition by SR proteins of splicing of a regulated adenovirus pre-mRNA. Nature. 1996;381(6582):535-538. doi:10.1038/381535a0. 34. Ule J, Stefani G, Meie A, et al. An RNA map predicting Nova-dependent splicing regulation. Nature. 2006;444(7119):580-586. doi:10.1038/nature05304. 35. Goren A, Ram O, Amit M , et al. Comparative analysis identifies exonic splicing regulatory sequences—The complex definition of enhancers and silencers. Mol Cell. 2006;22(6):769-781. doi:10.1016/j.molcel.2006.05.008. 36. Wan Y, Qu K, Zhang QC, et al. Landscape and variation of RNA secondary structure across the human transcriptome. Nature. 2014;505(7485):706-709. doi: 10.103 8/nature 12946. 37. Zhou H-L, Luo G, Wise JA, Lou H. Regulation of alternative splicing by local histone modifications: potential roles for RNA-guided mechanisms. Nucleic Acids Res. 2014;42(2):701 -713. doi: 10.1093/nar/gkt875. 38. Buratti E, Baralle FE. Influence of RNA secondary structure on the pre-mRNA splicing process. Mol Cell Biol. 2004;24(24): 10505-10514. doi: 10.1128/MCB.24.24.10505- 10514.2004. 39. Hiller M, Zhang Z, Backofen R, Stamm S. Pre-mRNA Secondary Structures Influence Exon Recognition. PLoS Genet. 2007;3(ll):e204. doi:10.1371/journal.pgen.0030204. 40. Shepard PJ, Choi E-A, Busch A, Hertel KJ. Efficient internal exon recognition depends on near equal contributions from the 3' and 5' splice sites. Nucleic Acids Res. 2011 ;39(20):8928-8937. doi: 10.1093/nar/gkr481. 41. Lareau LF, Inada M , Green RE, Wengrod JC, Brenner SE. Unproductive splicing of SR genes associated with highly conserved and ultraconserved D N A elements. Nature. 2007;446(7138):926-929. doi: 10.1038/nature05676. 42. Girard C, Will CL, Peng J, et al. Post-transcriptional spliceosomes are retained in nuclear speckles until splicing completion. Nat Commun. 2012;3:994. doi: 10.103 8/ncomms 1998. 43. Fong N , Kim H, Zhou Y, et al. Pre-mRNA splicing is facilitated by an optimal R N A polymerase II elongation rate. Genes Dev. 2014;28(23):2663-2676. doi:10.1101/gad.252106.114. 44. Jaillon O, Bouhouche K, Gout J-F, et al. Translational control of intron splicing in eukaryotes. Nature. 2008;451(7176):359-362. doi:10.1038/nature06495. 45. Fredericks A, Cygan K, Brown B, Fairbrother W. RNA-Binding Proteins: Splicing Factors and Disease. Biomolecules. 2015;5(2):893-909. doi:10.3390/biom5020893. 114 46. McKie AB, McHale JC, Keen TJ, et al. Mutations in the pre-mRNA splicing factor gene PRPC8 in autosomal dominant retinitis pigmentosa (RP13). Hum Mol Genet. 2001;10(15):1555-1562. 47. Vithana EN, Abu-Safieh L, Allen MJ, et al. A human homolog of yeast pre-mRNA splicing gene, PRP31, underlies autosomal dominant retinitis pigmentosa on chromosome 19ql3.4 (RPH). Mol Cell. 2001;8(2):375-381. 48. Chakarova CF, Hims M M , Bolz H, et al. Mutations in HPRP3, a third member of premRNA splicing factor genes, implicated in autosomal dominant retinitis pigmentosa. Hum Mol Genet. 2002;ll(l):87-92. 49. Arai T, Hasegawa M, Akiyama H, et al. TDP-43 is a component of ubiquitin-positive tau-negative inclusions in frontotemporal lobar degeneration and amyotrophic lateral sclerosis. Biochem Biophys Res Commun. 2006;351(3):602-611. doi:10.1016/j.bbrc.2006.10.093. 50. Sreedharan J, Blair LP, Tripathi VB, et al. TDP-43 Mutations in Familial and Sporadic Amyotrophic Lateral Sclerosis. Science. 2008;319(5870): 1668-1672. doi: 10.1126/science. 1154584. 51. Pellizzoni L, Kataoka N , Charroux B, Dreyfuss G. A novel function for SMN, the spinal muscular atrophy disease gene product, in pre-mRNA splicing. Cell. 1998;95(5):615-624. 52. Brauch K M , Karst ML, Herron KJ, et al. Mutations in Ribonucleic Acid Binding Protein Gene Cause Familial Dilated Cardiomyopathy. J Am Coll Cardiol. 2009;54(10):930-941. doi:10.1016/j.jacc.2009.05.038. 53. Sleeman J. Small nuclear RNAs and mRNAs: linking RNA processing and transport to spinal muscular atrophy. Biochem Soc Trans. 2013 ;41(4): 871-875. doi:10.1042/BST20120016. 54. Colombrita C, Onesto E, Tiloca C, Ticozzi N, Silani V, Ratti A. RNA-binding proteins and R N A metabolism: a new scenario in the pathogenesis of Amyotrophic lateral sclerosis. Arch Ital Biol. 2011; 149( 1): 83-99. 55. Merico D, Roifman M , Braunschweig U , et al. Compound heterozygous mutations in the noncoding RNU4ATAC cause Roifman Syndrome by disrupting minor intron splicing. Nat Commun. 2015;6:8718. doi:10.1038/ncomms9718. 56. Orr HT, Chung MY, Banfi S, et al. Expansion of an unstable trinucleotide CAG repeat in spinocerebellar ataxia type 1. Nat Genet. 1993;4(3):221-226. doi:10.1038/ng0793- 221. 57. Wang J, Pegoraro E, Menegazzo E, et al. Myotonic dystrophy: evidence for a possible dominant-negative RNA mutation. Hum Mol Genet. 1995;4(4):599-606. 58. Walsh MJ, Cooper-Knock J, Dodd JE, et al. Invited Review: Decoding the pathophysiological mechanisms that underlie RNA dysregulation in neurodegenerative disorders: a review of the current state of the art: Dysregulation of gene expression in 115 neurodegeneration. Neuropathol Appl Neurobiol. 2015;41(2): 109-134. doi:10.1111/nan.l2187. 59. Timchenko NA, Cai ZJ, Welm AL, Reddy S, Ashizawa T, Timchenko LT. RNA CUG repeats sequester CUGBP1 and alter protein levels and activity of CUGBP1. J Biol Chem. 2001;276(ll):7820-7826. doi:10.1074/jbc.M005960200. 60. White MC, Gao R, Xu W, et al. Inactivation of hnRNP K by Expanded Intronic AUUCU Repeat Induces Apoptosis Via Translocation of PKC8 to Mitochondria in Spinocerebellar Ataxia 10. Pearson CE, ed. PLoS Genet. 2010;6(6):el000984. doi: 10.1371/journal.pgen. 1000984. 61. Mykowska A, Sobczak K, Wojciechowska M , Kozlowski P, Krzyzosiak WJ. C A G repeats mimic CUG repeats in the misregulation of alternative splicing. Nucleic Acids Res. 2011;39(20):8938-8951. doi:10.1093/nar/gkr608. 62. Burset M, Seledtsov IA, Solovyev VV. Analysis of canonical and non-canonical splice sites in mammalian genomes. Nucleic Acids Res. 2000;28(21):4364-4375. 63. Houdayer C, Caux-Moncoutier V, Krieger S, et al. Guidelines for splicing analysis in molecular diagnosis derived from a set of 327 combined in silico/in vitro studies on BRCA1 and BRCA2 variants. Hum Mutat. 2012;33(8): 1228-1238. doi:10.1002/humu.22101. 64. Vorechovsky I. Aberrant 3' splice sites in human disease genes: mutation pattern, nucleotide structure and comparison of computational tools that predict their utilization. Nucleic Acids Res. 2006;34(16):4630-4641. doi:10.1093/nar/gkl535. 65. Buratti E, Chivers M , Královicova J, et al. Aberrant 5' splice sites in human disease genes: mutation pattern, nucleotide structure and comparison of computational tools that predict their utilization. Nucleic Acids Res. 2007;35(13):4250-4263. doi: 10.1093/nar/gkm402. 66. Královicova J, Lei H, Vořechovský I. Phenotypic consequences of branch point substitutions. HumMutat. 2006;27(8):803-813. doi:10.1002/humu.20362. 67. Gao K, Masuda A, Matsuura T, Ohno K. Human branch point consensus sequence is yUnAy. Nucleic Acids Res. 2008;36(7):2257-2267. doi:10.1093/nar/gkn073. 68. Sterne-Weiler T, Howard J, Mort M, Cooper DN, Sanford JR. Loss of exon identity is a common mechanism of human inherited disease. Genome Res. 2011;21(10): 1563-1571. doi:10.1101/gr.H8638.110. 69. Lim K H , Ferraris L, Filloux ME, Raphael BJ, Fairbrother WG. Using positional distribution to identify splicing elements and predict pre-mRNA processing defects in human genes. Proc Natl Acad Sci. 2011; 108(27): 11093-11098. doi:10.1073/pnas.H01135108. 70. Liu HX, Cartegni L, Zhang MQ, Krainer AR. A mechanism for exon skipping caused by nonsense or missense mutations in BRCA1 and other genes. Nat Genet. 2001;27(l):55-58. doi:10.1038/83762. 116 71. Gaildrat P, Krieger S, Di Giacomo D, et al. Multiple sequence variants of BRCA2 exon 7 alter splicing regulation. J Med Genet. 2012;49(10):609-617. doi: 10.1136/jmedgenet- 2012-100965. 72. Soukarieh O, Gaildrat P, Hamieh M , et al. Exonic Splicing Mutations Are More Prevalent than Currently Estimated and Can Be Predicted by Using In Silico Tools. Aretz S, ed. PLOS Genet. 2016;12(l):el005756. doi: 10.1371/journal.pgen. 1005756. 73. Disset A, Bourgeois CF, Benmalek N, Claustres M, Stevenin J, Tuffery-Giraud S. An exon skipping-associated nonsense mutation in the dystrophin gene uncovers a complex interplay between multiple antagonistic splicing elements. Hum Mol Genet. 2006; 15(6):999-1013. doi: 10.1093/hmg/ddlO 15. 74. Aznarez I, Zielenski J, Rommens JM, Blencowe BJ, Tsui L-C. Exon skipping through the creation of a putative exonic splicing silencer as a consequence of the cystic fibrosis mutation R553X. J Med Genet. 2007;44(5):341-346. doi:10.1136/jmg.2006.045880. 75. Goina E, Skoko N , Pagani F. Binding of DAZAP1 and hnRNPAl/A2 to an exonic splicing silencer in a natural BRCA1 exon 18 mutant. Mol Cell Biol. 2008;28(11):3850- 3860. doi: 10.1128/MCB.02253-07. 76. Dobrowolski SF, Andersen HS, Doktor TK, Andresen BS. The phenylalanine hydroxylase c.30C>G synonymous variation (p.GlOG) creates a common exonic splicing silencer. Mol Genet Metab. 2010;100(4):316-323. doi:10.1016/j.ymgme.2010.04.002. 77. Hernändez-Imaz E, Martin Y, de Conti L, et al. Functional Analysis of Mutations in Exon 9 of NF1 Reveals the Presence of Several Elements Regulating Splicing. PloS One. 2015;10(10):e0141735. doi:10.1371/journal.pone.0141735. 78. Cartegni L, Krainer AR. Disruption of an SF2/ASF-dependent exonic splicing enhancer in SMN2 causes spinal muscular atrophy in the absence of SMN1. Nat Genet. 2002;30(4):377-384. doi:10.1038/ng854. 79. Pedrotti S, Bielli P, Paronetto MP, et al. The splicing regulator Sam68 binds to a novel exonic splicing silencer and functions in SMN2 alternative splicing in spinal muscular atrophy. EMBOJ. 2010;29(7): 1235-1247. doi:10.1038/emboj.2010.19. 80. Lorson CL, Hahnen E, Androphy EJ, Wirth B. A single nucleotide in the SMN gene regulates splicing and is responsible for spinal muscular atrophy. Proc Natl Acad Sei U SA. 1999;96(11):6307-6311. 81. Monani UR, Lorson CL, Parsons DW, et al. A single nucleotide difference that alters splicing patterns distinguishes the SMA gene SMN1 from the copy gene SMN2. Hum MolGenet. 1999;8(7): 1177-1183. 82. Skoko N , Baralle M , Buratti E, Baralle FE. The pathological splicing mutation c.6792C>G in NF1 exon 37 causes a change of tenancy between antagonistic splicing factors. FEBSLett. 2008;582(15):2231-2236. doi: 10.1016/j.febslet.2008.05.018. 83. Woolfe A, Mullikin JC, Elnitski L. Genomic features defining exonic variants that 117 modulate splicing. Genome Biol. 2010;11(2):R20. doi:10.1186/gb-2010-ll-2-r20. 84. Lynch KW, Weiss A. A CD45 polymorphism associated with multiple sclerosis disrupts an exonic splicing silencer. J Biol Chem. 2001;276(26):24341-24347. doi:10.1074/jbc.M102175200. 85. Hobson GM, Huang Z, Sperle K, Stabley DL, Marks HG, Cambi F. A PLP splicing abnormality is associated with an unusual presentation of PMD. Ann Neurol. 2002;52(4):477-488. doi:10.1002/ana. 10320. 86. Wang E, Dimova N , Sperle K, et al. Deletion of a splicing enhancer disrupts PLP 1/DM20 ratio and myelin stability. Exp Neurol. 2008;214(2):322-330. doi: 10.1016/j.expneurol .2008.09.001. 87. Homolova K, Zavadakova P, Doktor TK, Schroeder LD, Kozich V, Andresen BS. The deep intronic c.903+469T>C mutation in the MTRR gene creates an SF2/ASF binding exonic splicing enhancer, which leads to pseudoexon activation and causes the cblE type of homocystinuria. HumMutat. 2010;31(4):437-444. doi:10.1002/humu.21206. 88. Masuda A, Shen X - M , Ito M , Matsuura T, Engel AG, Ohno K. hnRNP H enhances skipping of a nonfunctional exon P3A in CHRNA1 and a mutation disrupting its binding causes congenital myasthenic syndrome. Hum Mol Genet. 2008;17(24):4022- 4035. doi:10.1093/hmg/ddn305. 89. Kashima T, Rao N, Manley JL. An intronic element contributes to splicing repression in spinal muscular atrophy. Proc Natl Acad Sci USA. 2007;104(9):3426-3431. doi: 10.1073/pnas.0700343104. 90. Haque A, Buratti E, Baralle FE. Functional properties and evolutionary splicing constraints on a composite exonic regulatory element of splicing in CFTR exon 12. Nucleic Acids Res. 2010;38(2):647-659. doi :10.1093/nar/gkp 1040. 91. Tammaro C, Raponi M, Wilson D, Baralle D. BRCA1 Exon 11, a CERES (Composite Regulatory Element of Splicing) Element Involved in Splice Regulation. IntJMol Sci. 2014;15(7):13045-13059. doi:10.3390/ijmsl50713045. 92. Bruce SR, Peterson ML. Multiple features contribute to efficient constitutive splicing of an unusually large exon. Nucleic Acids Res. 2001;29(ll):2292-2302. 93. Chen LL, Sabripour M , Wu EF, Prieto VG, Fuller GN, Frazier ML. A mutation-created novel intra-exonic pre-mRNA splice site causes constitutive activation of KIT in human gastrointestinal stromal tumors. Oncogene. 2005;24(26):4271-4280. doi: 10.1038/sj. one. 1208587. 94. Salato VK, Rediske NW, Zhang C, Hastings M L , Munroe SH. An exonic splicing enhancer within a bidirectional coding sequence regulates alternative splicing of an antisense mRNA. UNA Biol. 2010;7(2): 179-190. 95. Nielsen KB, Sorensen S, Cartegni L, et al. Seemingly neutral polymorphic variants may confer immunity to splicing-inactivating mutations: a synonymous SNP in exon 5 of M C A D protects from deleterious mutations in a flanking exonic splicing enhancer. 118 Am J Hum Genet. 2007;80(3):416-432. doi: 10.1086/511992. 96. Houdayer C, Dehainault C, Mattler C, et al. Evaluation of in silico splice tools for decision-making in molecular diagnosis. Hum Mutat. 2008;29(7):975-982. doi:10.1002/humu.20765. 97. Lastella P, Surdo NC, Resta N , Guanti G, Stella A. In silico and in vivo splicing analysis of MLH1 and MSH2 missense mutations shows exon- and tissue-specific effects. BMC Genomics. 2006;7:243. doi:10.1186/1471-2164-7-243. 98. Flanigan K M , Dunn DM, von Niederhausern A, et al. Nonsense mutation-associated Becker muscular dystrophy: interplay between exon definition and splicing regulatory elements within the D M D gene. Hum Mutat. 2011;32(3):299-308. doi:10.1002/humu.21426. 99. Aissat A, de Becdeliěvre A, Golmard L, et al. Combined Computational-Experimental Analyses of CFTR Exon Strength Uncover Predictability of Exon-Skipping Level. HumMutat. 2013;34(6):873-881. doi:10.1002/humu.22300. 100. Varani L, Hasegawa M , Spillantini M G , et al. Structure of tau exon 10 splicing regulatory element RNA and destabilization by mutations of frontotemporal dementia and parkinsonism linked to chromosome 17. Proc Natl Acad Sci U S A. 1999;96(14):8229-8234. 101. Hasegawa M , Smith MJ, Iijima M, Tabira T, Goedert M . FTDP-17 mutations N279K and S305N in tau produce increased splicing of exon 10. FEBS Lett. 1999;443(2):93- 96. 102. Krawczak M , Reiss J, Cooper DN. The mutational spectrum of single base-pair substitutions in mRNA splice junctions of human genes: causes and consequences. Hum Genet. 1992;90(l-2):41-54. 103. Nakai K, Sakamoto H. Construction of a novel database containing aberrant splicing mutations of mammalian genes. Gene. 1994; 141 (2): 171 -177. 104. Roca X, Sachidanandam R, Krainer AR. Intrinsic differences between authentic and cryptic 5' splice sites. Nucleic Acids Res. 2003;31(21):6321-6333. 105. Haj Khelil A, Deguillien M, Moriniěre M, Ben Chibani J, Baklouti F. Cryptic splicing sites are differentially utilized in vivo. FEBS J. 2008;275(6):1150-1162. doi: 10.1111/j. 1742-4658.2008.06276.x. 106. Královicova J, Vorechovsky I. Global control of aberrant splice-site activation by auxiliary splicing sequences: evidence for a gradient in exon and intron definition. Nucleic Acids Res. 2007;35(19):6399-6413. doi:10.1093/nar/gkm680. 107. Huang CH, Reid M , Daniels G, Blumenfeld OO. Alteration of splice site selection by an exon mutation in the human glycophorin A gene. J Biol Chem. 1993;268(34):25902- 25908. 108. Teraoka SN, Telatar M , Becker-Catania S, et al. Splicing defects in the ataxiatelangiectasia gene, ATM: underlying mutations and consequences. Am J Hum Genet. 119 1999;64(6): 1617-1631. doi: 10.1086/302418. 109. Baralle M, Skoko N, Knezevich A, et al. NF1 mRNA biogenesis: Effect of the genomic milieu in splicing regulation of the NF1 exon 37 region. FEBS Lett. 2006;580(18):4449-4456. doi:10.1016/j.febslet.2006.07.018. 110. Raponi M , Upadhyaya M , Baralle D. Functional splicing assay shows a pathogenic intronic mutation in neurofibromatosis type 1 (NF1) due to intronic sequence exonization. HumMutat. 2006;27(3):294-295. doi:10.1002/humu.9412. 111. Ogino W, Takeshima Y, Nishiyama A, et al. Mutation analysis of the ornithine transcarbamylase (OTC) gene in five Japanese OTC deficiency patients revealed two known and three novel mutations including a deep intronic mutation. Kobe J Med Sei. 2007;53(5):229-240. 112. De Klein A, Riegman PH, Bijlsma EK, et al. A G—>A transition creates a branch point sequence and activation of a cryptic exon, resulting in the hereditary disorder neurofibromatosis 2. Hum Mol Genet. 1998;7(3):393-398. 113. Yamaguchi H, Fujimoto T, Nakamura S, et al. Aberrant splicing of the milk fat globuleEGF factor 8 (MFG-E8) gene in human systemic lupus erythematosus. Eur J Immunol. 2010;40(6):1778-1785. doi:10.1002/eji.200940096. 114. Will K, Dörk T, Stuhrmann M, et al. A novel exon in the cystic fibrosis transmembrane conductance regulator gene activated by the nonsense mutation E92X in airway epithelial cells of patients with cystic fibrosis. J Clin Invest. 1994;93(4): 1852-1859. doi:10.1172/JCI117172. 115. Chen X, Truong T-TN, Weaver J, et al. Intronic alterations in BRCA1 and BRCA2 : effect on mRNA splicing fidelity and expression. Hum Mutat. 2006;27(5):427-435. doi:10.1002/humu.20319. 116. Hwang D-Y, Hung C-C, Riepe FG, et al. CYP17A1 Intron Mutation Causing Cryptic Splicing in 17a-Hydroxylase Deficiency. Planas JV, ed. PLoS ONE. 2011;6(9):e25492. doi:10.1371/journal.pone.0025492. 117. Savino M , Bordello A, d'Apolito M , et al. Spectrum ofFANCA mutations in Italian Fanconi anemia patients: Identification of six novel alleles and phenotypic characterization of the S858R variant. Hum Mutat. 2003;22(4):338-339. doi:10.1002/humu.9180. 118. Stucki M , Suormala T, Fowler B, Valle D, Baumgartner MR. Cryptic Exon Activation by Disruption of Exon Splice Enhancer: NOVEL MECHANISM CAUSING 3METHYLCROTONYL-CoA CARBOXYLASE DEFICIENCY. J Biol Chem. 2009;284(42):28953-28957. doi:10.1074/jbc.M109.050674. 119. Sorel N , Mayeur-Rousse C, Deverriere S, et al. Comprehensive characterization of a novel intronic pseudo-exon inserted within an el4/a2 BCR-ABL rearrangement in a patient with chronic myeloid leukemia. J Mol Diagn JMD. 2010;12(4):520-524. doi:10.2353/jmoldx.2010.090218. 120 120. Baralle D, Lucassen A, Buratti E. Missed threads. The impact of pre-mRNA splicing defects on clinical practice. EMBO Rep. 2009; 10(8): 810-816. doi:10.1038/embor.2009.170. 121. Thompson BA, Spurdle AB, Plazzer J-P, et al. Application of a 5-tiered scheme for standardized classification of 2,360 unique mismatch repair gene variants in the InSiGHT locus-specific database. Nat Genet. 2014;46(2):107-115. doi:10.1038/ng.2854. 122. van der Klift HM, Jansen AML, van der Steenstraten N , et al. Splicing analysis for exonic and intronic mismatch repair gene variants associated with Lynch syndrome confirms high concordance between minigene assays and patient RNA analyses. Mol Genet Genomic Med. 2015;3(4):327-345. doi:10.1002/mgg3.145. 123. de la Hoya M , Soukarieh O, Löpez-Perolio I, et al. Combined genetic and splicing analysis of BRCA1 c.[594-2A>C; 641A>G] highlights the relevance of naturally occurring in-frame transcripts for developing disease gene variant classification algorithms. HumMol Genet. March 2016:ddw094. doi:10.1093/hmg/ddw094. 124. Buratti E, Baralle M , Baralle FE. From single splicing events to thousands: the ambiguous step forward in splicing research. Brief Funct Genomics. 2013; 12(1):3-12. doi:10.1093/bfgp/els048. 125. Shapiro MB, Senapathy P. RNA splice junctions of different classes of eukaryotes: sequence statistics and functional implications in gene expression. Nucleic Acids Res. 1987;15(17):7155-7174. doi:10.1093/nar/15.17.7155. 126. Bürge C, Karlin S. Prediction of complete gene structures in human genomic DNA. J Mol Biol. 1997;268(l):78-94. doi: 10.1006/jmbi. 1997.0951. 127. Reese M G , Eeckman FH, Kulp D, Haussler D. Improved Splice Site Detection in Genie. JComputBiol. 1997;4(3):311-323. doi: 10.1089/cmb. 1997.4.311. 128. Freund M , Asang C, Kammler S, et al. A novel approach to describe a U l snRNA binding site. Nucleic Acids Res. 2003;31(23):6963-6975. 129. Sharma N , Sosnay PR, Ramalho AS, et al. Experimental Assessment of Splicing Variants Using Expression Minigenes and Comparison with In Silico Predictions. Hum Mutat. 2014;35(10): 1249-1259. doi:10.1002/humu.22624. 130. Doktor TK, Schroeder LD, Vested A, et al. SMN2 exon 7 splicing is inhibited by binding of hnRNP A l to a common ESS motif that spans the 3' splice site. Hum Mutat. 2011;32(2):220-230. doi:10.1002/humu.21419. 131. Desmet F-O, Hamroun D, Lalande M , Collod-Beroud G, Claustres M , Beroud C. Human Splicing Finder: an online bioinformatics tool to predict splicing signals. Nucleic Acids Res. 2009;37(9):e67-e67. doi:10.1093/nar/gkp215. 132. Zampieri S, Buratti E, Dominissini S, et al. Splicing mutations in glycogen-storage disease type II: evaluation of the full spectrum of mutations and their relation to patients' phenotypes. Eur J Hum Genet EJHG. 2011; 19(4) :422-431. 121 doi:10.1038/ejhg.2010.188. 133. Steffensen AY, Dandanell M , Jonson L, et al. Functional characterization of BRCA1 gene variants by mini-gene splicing assay. Eur J Hum Genet. 2014;22(12):1362-1368. doi:10.1038/ejhg.2014.40. 134. Jian X, Boerwinkle E, Liu X. In silico prediction of splice-altering single nucleotide variants in the human genome. Nucleic Acids Res. 2014;42(22): 13534-13544. doi:10.1093/nar/gkul206. 135. Jian X, Boerwinkle E, Liu X. In silico tools for splicing defect prediction: a survey from the viewpoint of end users. Genet Med Off J Am Coll Med Genet. 2014;16(7):497- 503. doi:10.1038/gim.2013.176. 136. Yeo G, Burge CB. Maximum Entropy Modeling of Short Sequence Motifs with Applications to RNA Splicing Signals. J Comput Biol. 2004;ll(2-3):377-394. doi:10.1089/1066527041410418. 137. Divina P, Kvitkovicova A, Buratti E, Vorechovsky I. Ab initio prediction of mutationinduced cryptic splice-site activation and exon skipping. Eur J Hum Genet. 2009;17(6):759-765. doi:10.1038/ejhg.2008.257. 138. Kol G, Lev-Maor G, Ast G. Human-mouse comparative analysis reveals that branchsite plasticity contributes to splicing regulation. Hum Mol Genet. 2005; 14(11): 1559- 1568. doi:10.1093/hmg/ddil64. 139. Plass M , Agirre E, Reyes D, Camara F, Eyras E. Co-evolution of the branch site and SR proteins in eukaryotes. Trends Genet TIG. 2008;24(12):590-594. doi: 10.1016/j.tig.2008.10.004. 140. Schwartz SH, Silva J, Burstein D, Pupko T, Eyras E, Ast G. Large-scale comparative analysis of splicing signals and their corresponding splicing factors in eukaryotes. Genome Res. 2008;18(1):88-103. doi:10.1101/gr.6818908. 141. Corvelo A, Hallegger M , Smith CWJ, Eyras E. Genome-Wide Association between Branch Point Properties and Alternative Splicing. Meyer IM, ed. PLoS Comput Biol. 2010;6(ll):el001016. doi:10.1371/journal.pcbi.l001016. 142. Liu HX, Zhang M , Krainer AR. Identification of functional exonic splicing enhancer motifs recognized by individual SR proteins. Genes Dev. 1998;12(13): 1998-2012. 143. Cartegni L. ESEfinder: a web resource to identify exonic splicing enhancers. Nucleic Acids Res. 2003;31(13):3568-3571. doi:10.1093/nar/gkg616. 144. Fairbrother WG. Predictive Identification of Exonic Splicing Enhancers in Human Genes. Science. 2002;297(5583):1007-1013. doi: 10.1126/science. 1073774. 145. Fairbrother WG, Yeo GW, Yeh R, et al. RESCUE-ESE identifies candidate exonic splicing enhancers in vertebrate exons. Nucleic Acids Res. 2004;32(Web Server issue):W187-W190. doi: 10.1093/nar/gkh393. 146. Sironi M , Menozzi G, Riva L, et al. Silencer elements as possible inhibitors of 122 pseudoexon splicing. Nucleic Acids Res. 2004;32(5): 1783-1791. doi:10.1093/nar/gkh341. 147. Zhang XH-F, Chasin L A . Computational definition of sequence motifs governing constitutive exon splicing. Genes Dev. 2004;18(11): 1241-1250. doi:10.1101/gad.H95304. 148. Yeo GW, Van Nostrand EL, Nostrand ELV, Liang TY. Discovery and analysis of evolutionarily conserved intronic splicing regulatory elements. PLoS Genet. 2007;3(5):e85. doi:10.1371/journal.pgen.0030085. 149. Voelker RB, Berglund JA. A comprehensive computational characterization of conserved mammalian intronic sequences reveals conserved motifs associated with constitutive and alternative splicing. Genome Res. 2007; 17(7): 1023-103 3. doi:10.1101/gr.6017807. 150. Zhang C, L i W-H, Krainer AR, Zhang MQ. RNA landscape of evolution for optimal exon and intron discrimination. Proc Natl Acad Sei USA. 2008;105(15):5797-5802. doi:10.1073/pnas.0801692105. 151. Erkelenz S, Theiss S, Otte M , Widera M , Peter JO, Schaal H. Genomic HEXploring allows landscaping of novel potential splicing regulatory elements. Nucleic Acids Res. 2014;42(16): 10681 -10697. doi: 10.1093/nar/gku736. 152. Wang Z, Rolish ME, Yeo G, Tung V, Mawson M , Bürge CB. Systematic identification and analysis of exonic splicing silencers. Cell. 2004; 119(6):831-845. doi:10.1016/j.cell.2004.11.010. 153. Ke S, Shang S, Kalachikov SM, et al. Quantitative evaluation of all hexamers as exonic splicing elements. Genome Res. 2011;21(8): 1360-1374. doi: 10.1101/gr. 119628.110. 154. Barash Y, Vaquero-Garcia J, Gonzalez-Vallinas J, et al. AVISPA: a web tool for the prediction and analysis of alternative splicing. Genome Biol. 2013;14(10):R114. doi:10.1186/gb-2013-14-10-rll4. 155. Xiong HY, Alipanahi B, Lee LJ, et al. The human splicing code reveals new insights into the genetic determinants of disease. Science. 2015;347(6218):1254806-1254806. doi: 10.1126/science. 1254806. 156. Piva F, Giulietti M , Burini AB, Principato G. SpliceAid 2: a database of human splicing factors expression data and RNA target motifs. Hum Mutat. 2012;33(l):81-85. doi:10.1002/humu.21609. 157. Raponi M , Kralovicova J, Copson E, et al. Prediction of single-nucleotide substitutions that result in exon skipping: identification of a splicing silencer in BRCA1 exon 6. HumMutat. 2011;32(4):436-444. doi:10.1002/humu.21458. 158. Schwartz S, Hall E, Ast G. SROOGLE: webserver for integrative, user-friendly visualization of splicing signals. Nucleic Acids Res. 2009;37(Web Server):W189W192. doi:10.1093/nar/gkp320. 159. Campos B, Diez O, Domenech M , et al. RNA analysis of eight BRCA1 and BRCA2 123 unclassified variants identified in breast/ovarian cancer families from Spain: MUTATIONS IN BRIEF. HumMutat. 2003;22(4):337-337. doi:10.1002/humu.9176. 160. Auclair J, Busine MP, Navarro C, et al. Systematic mRNA analysis for the effect ofMLHl andMSH2 missense and silent mutations on aberrant splicing. Hum Mutat. 2006;27(2): 145-154. doi:10.1002/humu.20280. 161. ChasinLA. Searching for splicing motifs. Adv Exp Med Biol. 2007;623:85-106. 162. Zatkova A, Messiaen L, Vandenbroucke I, et al. Disruption of exonic splicing enhancer elements is the principal cause of exon skipping associated with seven nonsense or missense alleles of NF 1. Hum Mutat. 2004;24(6):491-501. doi:10.1002/humu.20103. 163. Di Giacomo D, Gaildrat P, Abuli A, et al. Functional Analysis of a Large set of BRCA2 exon 7 Variants Highlights the Predictive Value of Hexamer Scores in Detecting Alterations of Exonic Splicing Regulatory Elements: H U M A N MUTATION. Hum Mutat. 2013;34(11):1547-1557. doi:10.1002/humu.22428. 164. Stadler MB, Shomron N , Yeo GW, Schneider A, Xiao X, Bürge CB. Inference of Splicing Regulatory Activities by Sequence Neighborhood Analysis. PLoS Genet. 2006;2(ll):el91. doi:10.1371/journal.pgen.0020191. 165. Baralle D. Splicing in action: assessing disease causing sequence changes. J Med Genet. 2005;42(10):737-748. doi:10.1136/jmg.2004.029538. 166. Corley M , Solem A, Qu K, Chang HY, Laederach A. Detecting riboSNitches with RNA folding algorithms: a genome-wide benchmark. Nucleic Acids Res. 2015;43(3):1859- 1868. doi:10.1093/nar/gkv010. 167. Kralovicova J, Patel A, Searle M , Vorechovsky I. The role of short R N A loops in recognition of a single-hairpin exon derived from a mammalian-wide interspersed repeat. RNA Biol. 2015;12(l):54-69. doi:10.1080/15476286.2015.1017207. 168. Patterson DJ, Yasuhara K, Ruzzo WL. Pre-mRNA secondary structure prediction aids splice site prediction. Pac Symp Biocomput Pac Symp Biocomput. 2002:223-234. 169. Marashi S-A, Eslahchi C, Pezeshk H, Sadeghi M . Impact of RNA structure on the prediction of donor and acceptor splice sites. BMC Bioinformatics. 2006;7':297. doi:10.1186/1471-2105-7-297. 170. Muro AF, Caputi M , Pariyarath R, Pagani F, Buratti E, Baralle FE. Regulation of fibronectin EDA exon alternative splicing: possible role of RNA secondary structure for enhancer display. Mol Cell Biol. 1999;19(4):2657-2671. 171. Donahue CP, Muratore C, Wu JY, Kosik KS, Wolfe MS. Stabilization of the tau exon 10 stem loop alters pre-mRNA splicing. J Biol Chem. 2006;281(33):23302-23306. doi:10.1074/jbc.C600143200. 172. Zuker M . Mfold web server for nucleic acid folding and hybridization prediction. Nucleic Acids Res. 2003;31(13):3406-3415. 173. Knudsen B, Hein J. Pfold: R N A secondary structure prediction using stochastic 124 context-free grammars. Nucleic Acids Res. 2003;31(13):3423-3428. 174. Lorenz R, Bernhart SH, Höner Zu Siederdissen C, et al. ViennaRNA Package 2.0. Algorithms Mol Biol AMB. 2011;6:26. doi:10.1186/1748-7188-6-26. 175. Sabarinathan R, Tafer H, Seemann SE, Hofacker IL, Stadler PF, Gorodkin J. The RNAsnp web server: predicting SNP effects on local RNA secondary structure. Nucleic Acids Res. 2013;41(Web Server issue):W475-W479. doi:10.1093/nar/gkt291. 176. Salari R, Kimchi-Sarfaty C, Gottesman M M , Przytycka TM. Sensitive measurement of single-nucleotide polymorphism-induced changes of RNA conformation: application to disease studies. Nucleic Acids Res. 2013;41(l):44-53. doi:10.1093/nar/gksl009. 177. Halvorsen M , Martin JS, Broadaway S, Laederach A. Disease-associated mutations that alter the RNA structural ensemble. PLoS Genet. 2010;6(8):el001074. doi: 10.1371/journal.pgen. 1001074. 178. Cooper TA. Use of minigene systems to dissect alternative splicing elements. Methods. 2005;37(4):331-340. doi:10.1016/j.ymeth.2005.07.015. 179. Naruse H, Ikawa N , Yamaguchi K, et al. Determination of splice-site mutations in Lynch syndrome (hereditary non-polyposis colorectal cancer) patients using functional splicing assay. Fam Cancer. 2009;8(4):509-517. doi: 10.1007/s 10689-009-9280-6. 180. Grodeckä L, Lockerovä P, Ravcukovä B, et al. Exon First Nucleotide Mutations in Splicing: Evaluation of In Silico Prediction Tools. Spilianakis CB, ed. PLoS ONE. 2014;9(2):e89570. doi:10.1371/journal.pone.0089570. 181. Tournier I, Vezain M , Martins A, et al. A large fraction of unclassified variants of the mismatch repair genes MLH1 and MSH2 is associated with splicing defects. Hum Mutat. 2008;29(12): 1412-1424. doi:10.1002/humu.20796. 182. Stamm S, Smith CWJ, Lührmann R, eds. Alternative Pre-mRNA Splicing: Theory and Protocols. Weinheim: Wiley-Blackwell; 2012. 183. Spurdle AB, Couch FJ, Hogervorst FBL, Radice P, Sinilnikova OM, for the IARC Unclassified Genetic Variants Working Group. Prediction and assessment of splicing alterations: implications for clinical testing. Hum Mutat. 2008;29(11): 1304-1313. doi:10.1002/humu.20901. 184. Walker LC, Whiley PJ, Houdayer C, et al. Evaluation of a 5-Tier Scheme Proposed for Classification of Sequence Variants Using Bioinformatic and Splicing Assay Data: Inter-Reviewer Variability and Promotion of Minimum Reporting Guidelines. Hum Mutat. 2013;34(10): 1424-1431. doi:10.1002/humu.22388. 185. Plön SE, Eccles DM, Easton D, et al. Sequence variant classification and reporting: recommendations for improving the interpretation of cancer susceptibility genetic test results. Hum Mutat. 2008;29(11): 1282-1291. doi:10.1002/humu.20880. 186. van der Klift HM, Jansen AML, van der Steenstraten N , et al. Splicing analysis for exonic and intronic mismatch repair gene variants associated with Lynch syndrome confirms high concordance between minigene assays and patient RNA analyses. Mol 125 Genet Genomic Med. 2015;3(4):327-345. doi:10.1002/mgg3.145. 187. Whiley PJ, de la Hoya M, Thomassen M, et al. Comparison of mRNA Splicing Assay Protocols across Multiple Laboratories: Recommendations for Best Practice in Standardized Clinical Testing. Clin Chem. 2014;60(2):341-352. doi: 10.13 73/clinchem.2013.210658. 188. Picard C, Al-Herz W, Bousfiha A, et al. Primary Immunodeficiency Diseases: an Update on the Classification from the International Union of Immunological Societies Expert Committee for Primary Immunodeficiency 2015. J Clin Immunol. 2015;35(8):696-726. doi:10.1007/sl0875-015-0201-l. 189. de Vries E, Clinical Working Party of the European Society for Immunodeficiencies (ESID). Patient-centred screening for primary immunodeficiency: a multi-stage diagnostic protocol designed for non-immunologists. Clin Exp Immunol. 2006;145(2):204-214. doi:10.1111/j.l365-2249.2006.03138.x. 190. O'Keefe AW, Haibrich M , Ben-Shoshan M , McCusker C. Primary immunodeficiency for the primary care provider. Paediatr Child Health. 2016;21(2):el0-el4. 191. Abbas A K , Lichtman A H , Pillai S. Cellular and Molecular Immunology. Eighth edition. Philadelphia, PA: Elsevier Saunders; 2015. 192. Du L, Gatti RA. Potential therapeutic applications of antisense morpholino oligonucleotides in modulation of splicing in primary immunodeficiency diseases. J ImmunolMethods. 2011;365(l-2):l-7. doi:10.1016/j.jim.2010.12.001. 193. Santisteban I, Arredondo-Vega FX, Kelly S, et al. Novel splicing, missense, and deletion mutations in seven adenosine deaminase-deficient patients with late/delayed onset of combined immunodeficiency disease. Contribution of genotype to phenotype. JClinlnvest. 1993;92(5):2291-2302. doi:10.1172/JCI116833. 194. Markert M L , Finkel BD, McLaughlin TM, et al. Mutations in purine nucleoside Phosphorylase deficiency. Hum Mutat. 1997;9(2): 118-121. doi: 10.1002/(SICI)l098- 1004(1997)9:2<118::AID-HUMU3>3.0.CO;2-5. 195. Arredondo-Vega FX. Adenosine deaminase deficiency with mosaicism for a "secondsite suppressor" of a splicing mutation: decline in revertant T lymphocytes during enzyme replacement therapy. Blood. 2002;99(3): 1005-1013. doi :10.1182/blood.V99.3.1005. 196. Prod'homme T, Dekel B, Barbieri G, et al. Splicing defect in RFXANK results in a moderate combined immunodeficiency and long-duration clinical course. Immunogenetics. 2003;55(8):530-539. doi: 10.1007/s00251-003-0609-2. 197. Wada T, Yasui M , Torna T, et al. Detection of T lymphocytes with a second-site mutation in skin lesions of atypical X-linked severe combined immunodeficiency mimicking Omenn syndrome. Blood. 2008;112(5):1872-1875. doi: 10.1182/blood-2008- 04-149708. 198. Chapgier A, Kong X-F, Boisson-Dupuis S, et al. A partial form of recessive STAT1 126 deficiency in humans. J Clin Invest. 2009;119(6):1502-1514. doi:10.1172/JCI37083. 199. Gil J, Busto EM, Garcillän B, et al. A leaky mutation in CD3D differentially affects aß and y8 T cells and leads to a Taß-Ty8+B+NK+ human SCID. J Clin Invest. 2011;121(10):3872-3876. doi:10.1172/JCI44254. 200. Somech R, Lev A, Grisaru-Soen G, Shiran SI, Simon AJ, Grunebaum E. Purine nucleoside Phosphorylase deficiency presenting as severe combined immune deficiency. Immunol Res. 2013;56(1):150-154. doi:10.1007/sl2026-012-8380-9. 201. 0rstavik KH, Kristiansen M , Knudsen GP, et al. Novel splicing mutation in the NEMO (LKK-gamma) gene with severe immunodeficiency and heterogeneity of X chromosome inactivation. Am J Med Genet A. 2006; 140(1):31-39. doi: 10.1002/aj mg.a. 31026. 202. Arredondo-Vega FX, Santisteban I, Kelly S, Schlossman CM, Umetsu DT, Hershfield MS. Correct splicing despite mutation of the invariant first nucleotide of a 5' splice site: a possible basis for disparate clinical phenotypes in siblings with adenosine deaminase deficiency. Am J Hum Genet. 1994;54(5):820-830. 203. Torgerson TR, Linane A, Moes N , et al. Severe food allergy as a variant of IPEX syndrome caused by a deletion in a noncoding region of the FOXP3 gene. Gastroenterology. 2007; 132(5): 1705-1717. doi:10.1053/j.gastro.2007.02.044. 204. Chapel H, Prevot J, Gaspar HB, et al. Primary Immune Deficiencies - Principles of Care. Front Immunol. 2014;5. doi:10.3389/fimmu.2014.00627. 205. Du L, Gatti RA. Progress toward therapy with antisense-mediated splicing modulation. Curr Opin Mol Ther. 2009; 11 (2): 116-123. 206. Havens MA, Duelli D M , Hastings ML. Targeting R N A splicing for disease therapy. Wiley Interdiscip Rev RNA. 2013;4(3):247-266. doi: 10.1002/wrna. 1158. 207. Wang D, Gao G. State-of-the-art human gene therapy: part II. Gene therapy strategies and clinical applications. DiscovMed. 2014; 18(98): 151-161. 208. Nakamura K, Du L, Tunuguntla R, et al. Functional characterization and targeted correction of ATM mutations identified in Japanese patients with ataxia-telangiectasia. HumMutat. 2012;33(1): 198-208. doi:10.1002/humu.21632. 209. Cavalieri S, Pozzi E, Gatti RA, Brusco A. Deep-intronic ATM mutation detected by genomic resequencing and corrected in vitro by antisense morpholino oligonucleotide (AMO). Eur J Hum Genet. 2013;21(7):774-778. doi:10.1038/ejhg.2012.266. 210. Bestas B, Turunen JJ, Blomberg K E M , et al. Splice-Correction Strategies for Treatment of X-Linked Agammaglobulinemia. Curr Allergy Asthma Rep. 2015; 15(3). doi:10.1007/sll882-014-0510-0. 211. Tahara M , Pergolizzi RG, Kobayashi H, et al. Trans-splicing repair of CD40 ligand deficiency results in naturally regulated correction of a mouse model of hyper-IgM X linked immunodeficiency. Nat Med. 2004;10(8):835-841. doi: 10.103 8/nm 1086. 127 212. Zayed H, Xia L, Yerich A, et al. Correction of DNA Protein Kinase Deficiency by Spliceosome-mediated RNA Trans-splicing and Sleeping Beauty Transposon Delivery. MolTher. 2007;15(7):1273-1279. doi:10.1038/sj.mt.6300178. 213. Berger A, Maire S, Gaillard M-C, Sahel J-A, Hantraye P, Bemelmans A-P. mRNA trans-splicing in gene therapy for genetic diseases. Wiley Interdiscip Rev RNA. 2016;7(4):487-498. doi:10.1002/wrna.l347. 214. Tajnik M, Rogalska ME, Bussani E, et al. Molecular Basis and Therapeutic Strategies to Rescue Factor IX Variants That Affect Splicing and Protein Function. PLoS Genet. 2016;12(5):el006082. doi: 10.1371/journal.pgen. 1006082. 215. Salewsky B, Hildebrand G, Rothe S, et al. Directed Alternative Splicing in Nijmegen Breakage Syndrome: Proof of Principle Concerning Its Therapeutical Application. Mol TherJAmSoc Gene Ther. 2016;24(1):117-124. doi:10.1038/mt.2015.144. 216. GuÄ©dard-MÄ©reuze SL, VachÄ© C, Molinari N , et al. Sequence contexts that determine the pathogenicity of base substitutions at position +3 of donor splice-sites. HumMutat. 2009;30(9): 1329-1339. doi:10.1002/humu.21070. 217. Fu Y, Masuda A, Ito M , Shinmi J, Ohno K. AG-dependent 3'-splice sites are predisposed to aberrant splicing due to a mutation at the first nucleotide of an exon. Nucleic Acids Res. 2011;39(10):4396-4404. doi:10.1093/nar/gkr026. 218. Wu S, Romfo CM, Nilsen TW, Green MR. Functional recognition of the 3' splice site A G by the splicing factor U2AF35. Nature. 1999;402(6763):832-835. doi:10.1038/45590. 219. Lockerovä, Pavla. Vliv substituentu pfi mutaci v prvnim nukleotidu exonü na sestfih pre-mRNA [online], http://is.muni.cz/th/356753/prif_m/. 220. Shariat N, Holladay CD, Cleary RK, Phillips JA, Patton JG. Isolated growth hormone deficiency type II caused by a point mutation that alters both splice site strength and splicing enhancer function. Clin Genet. 2008;74(6):539-545. doi: 10.1111/j. 1399- 0004.2008.01042.x. 221. Acedo A, Sanz DJ, Durän M, et al. Comprehensive splicing functional analysis of DNA variants of the BRCA2 gene by hybrid minigenes. Breast Cancer Res. 2012;14(3):R87. doi:10.1186/bcr3202. 222. Ars E, Serra E, Garcia J, et al. Mutations affecting mRNA splicing are the most common molecular defects in patients with neurofibromatosis type 1. Hum Mol Genet. 2000;9(2):237-247. 223. Lin Z, Thomas NJ, Wang Y, et al. Deletions within a CA-repeat-rich region of intron 4 of the human SP-B gene affect mRNA splicing. Biochem J. 2005;389(2):403-412. doi:10.1042/BJ20042032. 224. Putnam EA, Cho M , Zinn AB, Towbin JA, Byers PH, Milewicz DM. Delineation of the Marfan phenotype associated with mutations in exons 23-32 of the FBN1 gene. Am J Med Genet. 1996;62(3):233-242. doi:10.1002/(SICI)1096- 128 8628(19960329)62:3<233::AID-AJMG7>3.0.CO;2-U. 225. Faivre L, Collod-Beroud G, Callewaert B, et al. Clinical and mutation-type analysis from an international series of 198 probands with a pathogenic FBN1 exons 24-32 mutation. Eur J Hum Genet. 2009;17(4):491-501. doi:10.1038/ejhg.2008.207. 226. Maeda J, Kosaki K, Shiono J, Kouno K, Aeba R, Yamagishi H. Variable severity of cardiovascular phenotypes in patients with an early-onset form of Marfan syndrome harboring FBN1 mutations in exons 24-32. Heart Vessels. 2016;31(10): 1717-1723. doi: 10.1007/s003 80-016-0793-2. 227. More H, Humar B, Weber W, et al. Identification of seven novel germline mutations in the human E-cadherin (CDH1) gene. Hum Mutat. 2007;28(2):203. doi:10.1002/humu.9473. 228. Molinaro V, Pensotti V, Marabelli M , et al. Complementary molecular approaches reveal heterogeneous CDH1 germline defects in Italian patients with hereditary diffuse gastric cancer (HDGC) syndrome. Genes Chromosomes Cancer. 2014;53(5):432-445. doi:10.1002/gcc.22155. 129 6. Abbreviations 3'ss 3' splice site or acceptor splice site; the intron-exon border of the 3' end of an intron 5'ss 5' splice site or donor splice site; the exon-intron border of the 5' end of an intron A adenine ASO antisense oligonucleotide bp base pairs C cytosine CERES composite exonic regulatory element of splicing CLIP Cross-Linking and ImmunoPrecipitation DNA deoxyribonucleic acid E + 1 exons first nucleotides ESE exonic splicing enhancer ESS exonic splicing silencer G guanine hnRNP heterogeneous nuclear ribonucleoproteins ISE intronic splicing enhancer ISS intronic splicing silencer IncRNA long non coding RNA NGS next generation sequencing NMD nonsense mediated decay nt nucleotides PLD primary immunodeficiency PPS the longest uninterrupted polypyrimidine stretch in the PPT of the 3'ss 130 PPT polypyrimidine tract PTC premature termination codon RESCUE Relative Enhancer and Silencer Classification by Unanimous Enrichment RNA ribonucleic acid RNAseq high throughput RNA sequencing SCUD severe combined immunodeficiency SELEX systematic evolution of ligands by exponential enrichment SMA spinal muscular atrophy SMaRT spliceosomal-mediated RNA /raws-splicing SNP single nucleotide polymorphism snRNP small nuclear ribonucleoprotein SNV single nucleotide variant SRE splicing regulatory elements T thymine U2AF U2 auxiliary factor VUS variant of unknown significance 131 7. List of figures Figure 1 - Pre-mRNA splicing by the U2-type spliceosome 12 Figure 2 - Early complexes of the spliceosome assembly 13 Figure 3 - Modes of alternative splicing with its protein coding consequences 13 Figure 4 - Consensual sequences of splice sites and the branch site 17 Figure 5 - Basic categorization of splicing regulatory elements (SREs) 17 Figure 6 - Typical minigene preparation 36 132 8. List of tables Table 1 - Selected online tools for prediction of splice site changes 28 Table 2 - Selected tools predicting specific 3'ss features: branch point and PPT 29 Table 3 - Selected tools for prediction of SREs 31 Table 4 - Selected online engines combining multiple splicing-prediction tools 33 Table 5 - Selected online tools for prediction of RNA secondary structure and its changes 34 133 9. List of supplements Supplementary Notes to the submitted manuscript 'Systematic analysis of splicing defects in selected primary immunodeficiencies-related genes.' Titles and Legends to Supplementary Figures and Tables to the submitted manuscript 'Systematic analysis of splicing defects in selected primary immunodeficiencies-related genes.' Supplementary Figures to the submitted manuscript 'Systematic analysis of splicing defects in selected primary immunodeficiencies-related genes.' Supplementary Tables to the submitted manuscript 'Systematic analysis of splicing defects in selected primary immunodeficiencies-related genes.' 134 10. List of publications In thefield of splicing: 1. Grodecká L, Lockerová P, Ravčuková B, et al. Exon First Nucleotide Mutations in Splicing: Evaluation of In Silico Prediction Tools. Spilianakis CB, ed. PLoS ONE. 2014;9(2):e89570. doi:10.1371/journal.pone.0089570. 2. Šípek A, Grodecká L, Baxová A, et al. Novel FBN1 gene mutation and maternal germinal mosaicism as the cause of neonatal form of Marfan syndrome. Am J Med Genet A. 2014;164A(6):1559-1564. doi:10.1002/ajmg.a.36480. 3. Grodecká L, Kramářek M, Lockerová P, et al. No major effect of the CDH1 c.2440- 6 0 G mutation on splicing detected in last exon-specific splicing minigene assay. Genes Chromosomes Cancer. 2014;53(9):798-801. doi:10.1002/gcc.22186. 4. Grodecká L, Kováčova T, Kramářek M, et al. Detailed molecular characterization of a novel IDS exonic mutation associated with multiple pseudoexon activation. J Mol Med BerlGer. November 2016. doi:10.1007/s00109-016-1484-2. In thefield of functional relevance of PID-related variants: 5. Freiberger T, Ravcukovä B, Grodeckä L, et al. Sequence variants of the TNFRSF13B gene in Czech CVID and IgAD patients in the context of other populations. Hum Immunol. 2012;73(11): 1147-1154. doi:10.1016/j.humimm.2012.07.342. 6. Freiberger T, Grodeckä L, Ravcukovä B, et al. Association of FcRn expression with lung abnormalities and IVIG catabolism in patients with common variable immunodeficiency. Clin Immunol Orlando Fla. 2010;136(3):419-425. doi:10.1016/j.clim.2010.05.006. 7. Freiberger T, Ravcukovä B, Grodeckä L, et al. No association of FCRN promoter VNTR polymorphism with the rate of maternal-fetal IgG transfer. J Reprod Immunol. 2010;85(2):193-197. doi:10.1016/j.jri.2010.04.002. 135 11. Summary of the thesis main findings In the thesis, I have investigated effect of genes variants on the process of pre-mRNA splicing and tested the possibilities to assess this effect using in silico predictors and in vitro minigene splicing analysis. Together with my collaborators, we have inspected 111 individual genes variants in silico (mostly found in primary immunodeficiency (PID)-related genes), and 64 of them in vitro, using splicing minigene assay and RT-PCR. The data showed that 21.7 % of the tested variants in the PID-related genes affected splicing, mostly through a change of splice sites. Further, minigene analysis was proved reasonably reliable in the assessment of the existence/absence of splicing affection. The data also indicated that three algorithms (ESRseq, EX-SKIP and Hexplorer) may come useful for predictions on SRE-affection. Further, we have shown that polypyrimidine tract per se is not very reliable feature for assessment of splice site aberration upon mutations of the exons first nucleotide and we have proposed specific cut-off values that could help discerning such splicing-affecting from non-affecting variants. In addition, we have proposed possible novel mechanisms of splicing disruption by these variants. Finally, we have described a new mechanism of pseudoexon activation resulting from competition between de novo and authentic splice site. Author: Supervisor: Mgr. Lucie Grodecka doc. MUDr. Tomas Freiberger, PhD. 136