Tuesday, May 16, 2017

Lecture 19: Jennifer Lu (Steven Salzberg Lab)

KrakEN and Bracken.

Metagenomics is a rapidly growing field of study, driven in part by our ability to generate enormous amounts of DNA sequence rapidly and inexpensively. Since the human genome was first published in 2001 (The International Human Genome Sequencing Consortium, 2001; Venter et al., 2001), sequencing technology has become approximately one million times faster and cheaper, making it possible for individual labs to generate as much sequence data as the entire Human Genome Project in just a few days. In the context of metagenomics experiments, this makes it possible to sample a complex mixture of microbes by “shotgun” sequencing, which involves simply isolating DNA, preparing the DNA for sequencing, and sequencing the mixture as deeply as possible. Shotgun sequencing is relatively unbiased compared to targeted sequencing methods (Venter et al., 2004), including widely-used 16S ribosomal RNA sequencing, and it has the additional advantage that it captures any species with a DNA-based genome, including eukaryotes that lack a 16S rRNA gene. Because it is unbiased, shotgun sequencing can also be used to estimate the abundance of each taxon (species, genus, phylum, etc.) in the original sample, by counting the number of reads belonging to each taxon. 
Along with the technological advances, the number of finished and draft genomes has also grown exponentially over the past decade. At present, there are thousands of complete bacterial genomes, 20,000 draft bacterial genomes, and 80,000 full or partial virus genomes in the public GenBank archive (Benson et al., 2015). This rich resource of sequenced genomes now makes it possible to sequence uncultured, unprocessed microbial DNA from almost any environment, ranging from soil to the deep ocean to the human body, and use computational sequence comparisons to identify many of the formerly hidden species in these environments (Riesenfeld, Schloss & Handelsman, 2004). Several accurate methods have appeared that can align a sequence “read” to a database of microbial genomes rapidly and accurately (see below), but this step alone is not sufficient to estimate how much of a species is present. Complications arise when closely related species are present in the same sample–a situation that arises quite frequently–because many reads align equally well to more than one species. This requires a separate abundance estimation algorithm to resolve. In their recent article, Jennifer Lu and Steven Salzberg from Johns Hopkins University and their colleagues describe a new method, Bracken, that goes beyond simply classifying individual reads and computes the abundance of species, genera, or other taxonomic categories from the DNA sequences collected in a metagenomics experiment.
Number of reads within the Mycobacterium genus as assigned by Kraken (blue), estimated by Bracken (purple) and compared to the true read counts (green)[1].
Bracken (Bayesian Reestimation of Abundance after Classification with KrakEN) uses the taxonomic assignments made by Kraken, a very fast read-level classifier, along with information about the genomes themselves to estimate abundance at the species level, the genus level, or above. The authors of the study demonstrate that Bracken can produce accurate species- and genus-level abundance estimates even when a sample contains multiple near-identical species.

References:
[1] Lu J, Breitwieser FP, Thielen P, Salzberg SL. Bracken: Estimating species abundance in metagenomics data. PeerJ Computer Science. 2017 Jan 2;3:e104.
[2] Wood DE, Salzberg SL. Kraken: ultrafast metagenomic sequence classification using exact alignments. Genome biology. 2014 Mar 3;15(3):R46.
___________________________________________________________________________
Jennifer Lu is a Biomedical Engineering Ph.D. Candidate in Professor Steven Salzberg's lab at the Center for Computational Biology at Johns Hopkins University. With a background in Chemical and Biomolecular Engineering and Computer Science, Jennifer began her Ph.D. with the intent of applying her knowledge in Computer Science to Biomedical research. Currently, her research is focused on computational genomics and the usage of next-generation sequencing for diagnosing bacterial, fungal, or viral infections relating to human health and diseases. As part of this research, she develops and uses various computational methods for quantifying DNA sequence similarities and analyzing the genomes of human pathogens.

Tuesday, April 11, 2017

Lecture 18: Jeliazko Jeliazkov (Jeff Gray Lab)

Computational Modeling and Docking of Antibody Structures.

The vertebrate adaptive immune system is capable of promoting cells to degranulate or phagocytose nearly any foreign pathogen by producing immunoglobulin G (IgG) proteins (antibodies) that recognize a specific region (epitope) of a pathogenic molecule (antigen). The ability to bind diverse antigens requires a diverse population of antibodies, which is achieved through complex processes in bone marrow and lymphatic tissues, namely V(D)J recombination and somatic hypermutation. The diversity of antibodies is astonishing; the size of the theoretical naive antibody repertoire is estimated to be >1e13 in humans. In addition to their biological importance, antibodies are routinely used in biotechnology as probes and diagnostics, and dozens of antibodies have been approved as therapeutics.

Therapeutic monoclonal antibodies are a genre of biopharmaceuticals which has benefitted healthcare in various fields from oncology to immune and inflammatory disorders. Development of successful novel therapeutic antibodies requires an understanding of drug and disease mechanisms and the ability to stabilize, affinity mature, and humanize antibodies. Antibody structures can help overcome these challenges by providing atomic-level insights into structure–function relationships and the antibody–antigen interaction [e.g. see refs. (1–4)]. However, experimental techniques for obtaining antibody structures, like X-ray crystallography and nuclear magnetic resonance, are laborious, time-consuming and costly. Computational antibody structure prediction provides a fast and inexpensive route to obtain structures, including those which are not obtainable otherwise.

A schematic[1] of the modeling protocols (full flowcharts
for Rosetta Antibody and Rosetta SnugDock
are available in the original publications).
In their recent Nature Protocols article, Jeliazko Jeliazkov and Jeff Gray from Johns Hopkins Department of Chemical and Biomolecular Engineering and their international collaborators describe Rosetta-based computational protocols for predicting the 3D structure of an antibody from the sequence (RosettaAntibody) and then docking the antibody to protein antigens (SnugDock). Antibody modeling leverages canonical loop conformations to graft large segments from experimentally determined structures, as well as offering (i) energetic calculations to minimize loops, (ii) docking methodology to refine the VL–VH relative orientation and (iii) de novo prediction of the elusive complementarity determining region (CDR) H3 loop. To alleviate model uncertainty, antibody–antigen docking resamples CDR loop conformations and can use multiple models to represent an ensemble of conformations for the antibody, the antigen or both. These protocols can be run fully automated via the ROSIE web server or manually on a computer with user control of individual steps. For best results, the protocol requires roughly 1,000 CPU-hours for antibody modeling and 250 CPU-hours for antibody–antigen docking. Tasks can be completed in under a day by using public supercomputers. In the figure, the structure on the left shows the FV antibody domains predicted by homology modeling (heavy chain in dark blue with CDR H1 and H2 loops in orange and CDR H3 loop in red; light chain in yellow with its CDR loops in light blue). The structure on the right depicts an antibody–antigen structure output by docking (antigen in green).

References:
[1]Weitzner BD, Jeliazkov JR, Lyskov S, Marze N, Kuroda D, Frick R, Adolf-Bryfogle J, Biswas N, Dunbrack Jr RL, Gray JJ. Modeling and docking of antibody structures with Rosetta. Nature Protocols. 2017 Feb 1;12(2):401-16.
[2]Sircar A, Kim ET, Gray JJ. RosettaAntibody: antibody variable region homology modeling server. Nucleic acids research. 2009 May 20:gkp387.
______________________________________________________________________________
Jeliazko Jeliazkov received his B.S. in Physics from the University of Illinois at Urbana-Champaign. Since graduating, he joined the Program in Molecular Biophysics and is pursuing a Ph.D. under the tutelage of Prof. Jeffrey Gray. His research involves the computational prediction of protein–protein interactions, focusing in particular on antibody–antigen, disordered–ordered protein domain, and crystallographic (non-biological) protein–protein interactions.

Tuesday, April 4, 2017

Lecture 17: Prof. Sagar Khare

A New Weapon in the Fight Against Cancer and Viral Infections: Custom-Designed Enzymes
Arising out of natural selection, the structures of proteins (and their complexes with small molecules, nucleic acids, and other proteins) display exquisitely fine-tuned molecular recognition, which is critical for life to operate. Under selection conditions, accurate molecular recognition must be robust to random perturbations such as mutations. Yet, natural proteins are also evolvable — variation in a few amino acids can lead to profound changes in function, e.g. a new enzymatic activity can arise in an “old” enzyme. In other words, these molecular interactions have the fascinating property of being simultaneously functionally robust and plastic.
Fitness scoring function [1].
Characterizing the substrate specificity of protease enzymes is critical for illuminating the molecular basis of their diverse and complex roles in a wide array of biological processes. Rapid and accurate prediction of their extended substrate specificity would also aid in the design of custom proteases capable of selectively and controllably cleaving biotechnologically or therapeutically relevant targets. However, current in silico approaches for protease specificity prediction, rely on, and are therefore limited by, machine learning of sequence patterns in known experimental data. In his talk, Prof. Sagar Khare described a general approach for predicting peptidase substrates de novo using protein structure modeling and biophysical evaluation of enzyme–substrate complexes. His research team constructed atomic resolution models of thousands of candidate substrate–enzyme complexes for each of five model proteases belonging to the four major protease mechanistic classes—serine, cysteine, aspartyl, and metallo-proteases—and develop a discriminatory scoring function using enzyme design modules from Rosetta and AMBER's MMPBSA. They ranked putative substrates based on calculated interaction energy with a modeled near-attack conformation of the enzyme active site. Their results show that the energetic patterns obtained from these simulations can be used to robustly rank and classify known cleaved and uncleaved peptides and that these structural-energetic patterns have greater discriminatory power compared to purely sequence-based statistical inference. Combining sequence and energetic patterns using machine-learning algorithms further improves classification performance, and analysis of structural models provides physical insight into the structural basis for the observed specificities.
Summary of the CPG2 circular permutations [2].
In the second part of his talk, Prof. Khare talked about spatio-temporal design of enzymes. Carboxypeptidase G2 (CPG2) is an Food and Drug Administration (FDA)-approved enzyme drug used to treat methotrexate (MTX) toxicity in cancer patients receiving MTX treatment. It has also been used in directed enzyme-prodrug chemotherapy, but this strategy has been hampered by off-site activation of the prodrug by the circulating enzyme. The development of a tumor protease activatable CPG2, which could be achieved using a circular permutation of CPG2 fused to an inactivating ‘prodomain’, would aid in these applications. The research team reported the development of a protease accessibility-based screen to identify candidate sites for circular permutation in proximity of the CPG2 active site. The resulting six circular permutants showed similar expression, structure, thermal stability, and, in four cases, activity levels compared to the wild-type enzyme. They rationalize these results based on structural models of the permutants obtained using the Rosetta software by developing a cell growth-based selection system, and demonstrated that when fused to periplasm-directing signal peptides, one of the circular permutants confers MTX resistance in Escherichiacoli with equal efficiency as the wild-type enzyme. As the permutants have similar properties to wild-type CPG2, these enzymes are promising starting points for the development of autoinhibited, protease-activatable zymogen forms of CPG2 for use in therapeutic contexts[2]
Related References:
[1] Pethe, M. A., Rubenstein, A. B., & Khare, S. D. (2017). Large-scale Structure-based Prediction and Identification of Novel Protease Substrates using Computational Protein Design. Journal of Molecular Biology, 429(2), 220-236
[2] Yachnin, B. J., & Khare, S. D. (2017). Engineering carboxypeptidase G2 circular permutations for the design of an autoinhibited enzyme. Protein engineering, design & selection: PEDS, 1.
____________________________________________________________________________
Sagar Khare is an Assistant Professor at Rutgers University. He teaches Chemistry and Chemical Biology and seeks to understand the structural determinants of enzymatic specificity and reactivity using a combination of computational protein design and experimental characterization. His research team's goal is to develop a quantitative and predictive understanding of specificity at protein-ligand and protein-peptide interfaces; which will inform various therapeutic and synthetic applications. Prior to working at Rutgers, Prof. Sagar Khare was a Postdoctoral Fellow at the University of Washington in Seattle, Washington. He was also a Software Engineer for Affymax Research Institute in Bangalore, India. 

Tuesday, March 21, 2017

Lecture 16: Dr. Yukun Wang (Martin Ulmschneider Lab)

Experimentally guided simulations reveal the mechanism of action of antimicrobial peptides
Since their discovery over 100 years ago, thousands of pore-forming antimicrobial peptides (AMPs) have been identified, revealing great variations in size, secondary structure and sequence composition. Despite this wealth of data no common pore-forming motif has been discovered, and the molecular basis of antimicrobial activity remains poorly understood. One of the key difficulties has been the transience of pores formed by AMPs, which has proved challenging for experimental structure determination. In the absence of experimental structures a range of theoretical models have been proposed that try to rationalize how soluble AMPs, which carry multiple charged residues, can insert into hydrophobic membranes to form pores and what these pores might look like.
Spontaneous membrane insertion and intra-membrane
oligomerization of maculatin [1]
Scientists from the Johns Hopkins Institue of NanoBio technology reports[1] an experimentally guided unbiased simulation methodology that yields the mechanism of spontaneous pore assembly for the AMP maculatin at atomic resolution. Rather than a single pore, maculatin forms an ensemble of structurally diverse temporarily functional low-oligomeric pores, which mimic integral membrane protein channels in structure. These pores continuously form and dissociate in the membrane. Membrane permeabilization is dominated by hexa-, hepta- and octamers, which conduct water, ions and small dyes. Pores were shown to form by consecutive addition of individual helices to a transmembrane helix or helix bundle, in contrast to current poration models. The diversity of the pore architectures—formed by a single sequence—may be a key feature in preventing bacterial resistance and could explain why sequence–function relationships in AMPs remain elusive.
Related References:
[1] Wang, Y., Chen, C. H., Hu, D., Ulmschneider, M. B., & Ulmschneider, J. P. (2016). Spontaneous formation of structurally diverse membrane channel architectures from a single antimicrobial peptide. Nature Communications, 7.
[2] Wimley, W. C., & Hristova, K. (2011). Antimicrobial peptides: successes, challenges and unanswered questions. The Journal of membrane biology, 239(1-2), 27-34.
[3] Ulmschneider, M. B., Ulmschneider, J. P., Schiller, N., Wallace, B. A., Von Heijne, G., & White, S. H. (2014). Spontaneous transmembrane helix insertion thermodynamically mimics translocon-guided insertion. Nature communications, 5, 4863.
_______________________________________________________________________________
Dr. Yukun Wang did his undergrad in Biology in Nanjing Agricultural University. After graduation, he did his Ph.D. in Biophysics at the Institute of Natural Sciences & School Of Life Sciences And Biotechnology, Shanghai Jiao Tong University. Currently, he is a postdoctoral fellow at the Institute for Nano-Bio-Technology in the Whiting School of Engineering at The Johns Hopkins University working with Prof. Martin Ulmschneider. Dr. Wang's research interests span from the functional mechanism of antimicrobial peptides to developing free energy calculation algorithms for rare event dynamics in the lipid-peptides systems. He enjoys computational designing of membrane-active peptides and proteins.

Tuesday, February 28, 2017

Lecture 15: Prof. Alex MacKerell

From Lemkul, J. A. et. al., Chemical reviews, 2016 [2].
Recent developments in polarizable force field research
Atomistic molecular dynamics (MD) simulations have become an integral component of the tool set used to examine biomolecular systems. Since the first MD simulation of a protein in 1977 [1], the time scales and sizes of computationally tractable simulations have grown by orders of magnitude. Systems may now be simulated on the microsecond or millisecond time scale and contain over a million atoms due to increased computing power and ever-improving algorithms for parallel and graphical processing unit (GPU)-based computing. Such capabilities also allow for more rigorous testing of the accuracy of the models used for such MD simulations.
Molecular mechanics force fields that explicitly account for induced polarization represent the next generation of physical models for molecular dynamics simulations. Several methods exist for modeling induced polarization, and in this seminar, Prof. MacKerell reviewed the classical Drude oscillator model, in which electronic degrees of freedom are modeled by charged particles attached to the nuclei of their core atoms by harmonic springs. He described the latest developments in Drude force field parametrization and application, primarily in the last 15 years. Emphasis was placed on the Drude-2013 polarizable force field for proteins, DNA, lipids, and carbohydrates. He discussed its parametrization protocol, development history, and recent simulations of biologically interesting systems, highlighting specific studies in which induced polarization plays a critical role in reproducing experimental observables and understanding physical behavior. As the Drude oscillator model is computationally tractable and available in a wide range of simulation packages, it is anticipated that use of these more complex physical models will lead to new and important discoveries of the physical forces driving a range of chemical and biological phenomena.
Related References:
[1] McCammon, J. A., Gelin, B. R., & Karplus, M. (1977). Dynamics of folded proteins. Nature, 267(5612), 585.
[2] Lemkul, J. A., Huang, J., Roux, B., & MacKerell Jr, A. D. (2016). An empirical polarizable force field based on the classical drude oscillator model: development history and recent applications. Chemical reviews, 116(9), 4983-5013.
__________________________________________________________________________
Alex MacKerell received an A.S. in biology in 1979 from Gloucester County College, Sewell, NJ, followed by a B.S. in chemistry in 1981 from the University of Hawaii, Honolulu, HI, and a Ph.D. in biochemistry in 1985 from Rutgers University, New Brunswick, NJ. Subsequent training involved postdoctoral fellowships in the Department of Medical Biophysics, Karolinska Intitutet, Stockholm, Sweden, in the area of experimental and theoretical biophysics and in the Department of Chemistry, Harvard University in theoretical chemistry. Following one year as a visiting professor at Swarthmore College, Swarthmore, PA, he assumed his faculty position in the School of Pharmacy, University of Maryland, Baltimore in 1993. MacKerell is currently the Grollman-Glick Professor of Pharmaceutical Sciences in the School of Pharmacy and the Director of the University of Maryland Computer-Aided Drug Design Center. MacKerell is also cofounder and Chief Scientific Officer of SilcsBio LLC. Research interests include the development of theoretical chemistry methods, with emphasis on empirical force field development; structure–function studies of proteins, carbohydrates, and nucleic acids; and the application of theoretical methods to drug discovery.

Thursday, December 1, 2016

Lecture 14: David Holland (Margaret Johnson Lab)

Finding the right partner in a crowded world

Mammalian cells contain ∼21,000 genes encoding upwards to 100,000 protein types. 5-40% of cell volume is occupied by macromolecules, posing challenges for cell proteins to locate functional partners and increasing the risk of nonspecific (nonfunctional) interactions. Overexpressed proteins, in particular, will saturate functional partners, leaving leftovers for nonspecific binding instead. Eukaryotic cells have evolved various methods to help proteins function reliably, including compartmentalization, allostery, and structural and chemical properties of binding sites. Cells may optimize specificity as a function of their binding networks; first through concentration balance, next through network structure. Using the Gillespie algorithm, David Holland and Dr. Margaret Johnson from Johns Hopkins Biophysics simulated specific and nonspecific binding in 500 networks of 90-200 nodes with varying topological properties under equal, random, and stoichiometrically balanced protein concentrations. Binding affinities for all specific and nonspecific interactions were determined using a coarse-grained protein sequence model. The research team found out that the concentration balance significantly reduced the number of nonspecific interactions, as did local topological patterns (motifs) that allowed increased difference between specific and nonspecific binding affinities. 
Schematic illustrating (left) the binding of three proteins, where matching features indicate binding interfaces, (center) the corresponding interface-interaction network with orange and green indicating shared and non-shared binding interfaces, respectively, and (right) the protein-protein interaction network (from [1]).
To test if these motifs are selected for in real networks, they sampled possible interface-interaction networks (IINs) for the Clathrin-mediated endocytosis PPI network in yeast, using a Monte Carlo algorithm and a fitness function to select for features that allow increased specificity. The fitness function reproduced several features of the real IIN, including fragmentation, a scale-free degree distribution, the presence of square motifs and a low number of chain and triangle motifs. These features differed significantly from unbiased sampling. In conclusion, there is selective pressure on the evolution of real IINs to avoid nonfunctional interactions and that non-optimal features of the real IIN (chains, larger network components) likely serve a functional purpose.

Relevant Publications:
[1]Johnson, ME, & G Hummer (2013) “Evolutionary Pressure on the Topology of Protein Interface Interaction Networks.” J Phys Chem B 117:13098-106. 
[2]Johnson, ME, G Hummer (2013). “Interface-resolved network of protein-protein interactions.” PLoS Comput Biol 9(5):e1003065 
[3]Keil, C, E Verschueren, J Yang, & L Serrano (2013) “Integration of protein abundance and structure data reveals competition in the ErbB signaling network.” Science Signaling 6(306): ra109 [4]Johnson, ME, & G Hummer (2011) “Nonspecific binding limits the number of proteins in a cell and shapes their interaction networks.” PNAS 108(2):603-8. 
[5]Zhang, J, S Maslov, & EI Shakhnovich (2008) “Constraints imposed by non-functional protein-protein interactions on gene expression and proteome size.” Mol Sys Biol 4:210 
[6]Vavouri, T, JI Semple, R Garcia-Verdugo, & B Lehner (2009) “Intrinsic protein disorder and interaction promiscuity are widely associated with dosage sensitivity.” Cell 138: 198-208.
_________________________________________________________________________________

David Holland received his B.Sc. in Biomedical Engineering from the University of Virginia in 2011. After graduating, he joined the department of Biomedical Engineering at Johns Hopkins. He is currently earning his Ph.D. under Margaret Johnson where he studies the effects of protein abundance and protein-protein interaction network structure on protein mis-interactions. In his spare time, David practices taekwondo and is also teaching a course on network science. 

Tuesday, November 8, 2016

Lecture 13: Athena Chen (Margaret Johnson Lab)

Spatial Cell Modeling Methods 
Due to the complexity of cells, it is useful to use computational tools to understand and predict mechanisms of biological processes. Ordinary differential equations (ODEs) and partial differential equations (PDEs) have proven to be successful at modeling various large systems. However, to understand certain biological processes such as bacterial cell division, high spatial resolution is necessary to discern the underlying mechanisms and interactions. ODEs and PDEs, along with the Gillespie algorithm for stochastic modeling, do not provide spatial resolution at a single-particle level; they are all concentration-based methods. MCell, the Free Propagator Reweighting Algorithm (FPR), and Smoldyn are algorithms that provide single-particle resolution, but may be computationally expensive and may not necessarily yield accurate protein dynamics.

Change in concentration of molecule A in an irreversible 3D reaction A+A -> 0. In this parameter set, the initial distances between molecules causes an increased initial reaction rate.
In her recent work, Athena Chen and her mentors Dr. Margaret Johnson and Dr. Osman Yogurtcu in the Johns Hopkins Biophysics department, analyzed the strengths and limitations of ODEs, PDEs, Gillespie, MCell, Smoldyn, and FPR through establishing and performing a set of benchmark tests. Though all modeling methods extracted the correct equilibrium concentrations for most of the tested reactions, the resulting protein dynamics were not necessarily correct. Simulations at high rates and large densities showed that despite providing single-particle resolution, Smoldyn and MCell do not pick up single-particle effects where the distance between two molecules affects the probability of binding. Furthermore, the dynamics given by Smoldyn and MCell are dependent on the time step selected. On the other hand, FPR correctly identified single-particle effects and yielded dynamics independent of the selected timestep.
As an example of the effects of molecular geometry, diffusion, and stochasticity of protein dynamics, we examined a model for bacterial cell division. From oscillations in protein concentrations, the cell can identify the center of the cell to ensure identical offspring and division of genetic information. 
Written by Athena Chen
Relevant Articles:
1-Yogurtcu, Osman N., and Margaret E. Johnson. "Theory of bi-molecular association dynamics in 2D for accurate model and experimental parameterization of binding rates.The Journal of chemical physics 143.8 (2015): 084117.
2-Andrews, Steven S., et al. "Detailed simulations of cell biology with Smoldyn 2.1." PLoS Comput Biol 6.3 (2010): e1000705.
4-Kerr, Rex A., et al. "Fast Monte Carlo simulation methods for biological reaction-diffusion systems in solution and on surfaces." SIAM journal on scientific computing 30.6 (2008): 3126-3149.
_________________________________________________________________________________
Athena Chen is currently working on her Bachelors of Arts in Biophysics and Bachelors of Science in Applied Mathematics and Statistics from Johns Hopkins University. As part of Dr. Margaret Johnson’s lab in the department of Biophysics, she analyzes the accuracy of methods for modeling the dynamics of protein interactions. In her free time, she enjoys yoga and figure skating.