
September 28th 2011,
Congress Center Het Pand, Gent, Belgium
Bioinformatics: tools
in research

13:30 Welcome reception
14:00 Introduction: WOUD and bioinformatics (Prof. dr. Bruno
Verhasselt, Prof. dr. Jo Vandesompele)
14:05 Investigating the human gut flora using metagenomics (Prof. dr. Jeroen
Raes, VUB)
14:25 The post-genomic era: epigenetic sequencing
applications and data integration (Dr. ir.
Maté Ongenaert, CMGG - UGent)
14:45 Visualisation of large
datasets (Drs. Geert Trooskens, BIOBIX - UGent)
15:00 Race against the sequencing machine: processing of
raw DNA sequence data at the Genomics Core (Prof. dr. Luc Dehaspe, Genomics Core - UZLeuven)
15:20 Coffee break
15:45 Integrative transcriptomics
to study non-coding RNA functions (Dr. ir.
Pieter Mestdagh, CMGG - UGent)
16:05 Large scale machine learning challenges for systems
biology (Dr. Yvan Saeys, VIB - UGent)
16:25 Pathway analysis: example from the bench (Drs. Jolien Vermeire,
HIVlab, Department of Clinical Chemistry,
Microbiology and Immunology - UGent)
16:40 Proteomics and cross-omics
integration (Prof. dr. Lennart Maertens (VIB - UGent)
17:00 Closure
Prof. dr. Jeroen Raes
Bioinformatics and (eco-)systems
biology, Department of Molecular and Cellular Interactions, VIB - Vrije Universiteit Brussel
Investigating the human gut flora using metagenomics
Meta‐omics (metagenomics, metatranscriptomics, metaproteomics) are powerful tools for the analysis of the (unculturable fraction of) microbial communities. Because of its complexity, meta‐omics data has required the development of novel computational analysis tools to determine the functional and phylogenetic composition of the sampled community. However, to go from a metagenomic ‘parts list’ (i.e. a bag of genes) to an initial understanding of the ecosystem structure and functioning, current tools are not sufficient (Raes & Bork, Nat Rev Microbiol 2009). I will present a range of approaches to analyze metagenomes, interpret metabolic changes and identify biologically and clinically relevant features from meta‐omics data with specific application to the human microbiome (Arumugam*, Raes* et al. Nature 2011).
Enterotypes of the human gut microbiome.
(2011) Arumugam M, Raes J
et al NATURE, 473, 174-80
Molecular eco-systems biology: towards an
understanding of community function. (2008)
Raes J, Bork P NATURE REVIEWS MICROBIOLOGY, 6,
693-9
Dr. ir. Maté Ongenaert
Center for Medical Genetics,
Ghent University
The post-genomic era: epigenetic sequencing applications and data
integration
The past decade is known as the post-genomic era. Ever since the first published human genomes, the pace to determine new genomes ever increased. In addition, a number of new sequencing applications gave access to previously unexplored areas at a genome-wide scale such as whole epigenomes.
In this talk, the data generated from a number of sequencing techniques to determine whole DNA-methylomes and whole genome histone marks will be discussed.
Main goal: to convince scientists that the analysis tools have matured to a level that, using a good manual and insight in the mechanisms behind the analysis, they can do their own basic analyses.
Starting from a raw sequence file, over quality control to mapping to the reference genome, peak calling, visualization and identification of differentially methylated sites: within the time-frame of this talk, the entire process will be demonstrated.
As epigenetics regulates genomic processes and literally is a layer above genetics, able to fine-tune regulatory processes, several layers of information should be look at to understand the underlying mechanisms.
Important aspect in the analysis of epigenetic datasets thus is the integration of several data sources (expression results, re-expression results, DNA-methylation information and histone-modifications).
Drs. Geert Trooskens
BIOBIX, UGent
Visualisation of large datasets
To be completed
Prof. dr. Luc Dehaspe
Bioinformatician, Genomics Core, UZ Leuven
Race against the sequencing machine: processing of raw DNA sequence
data at the Genomics Core
To grow and function, a living organism unconscious and continuous reads instructions from the DNA sequence in each of its cells. Thanks to the advances in DNA sequencing technologies, scientists are increasingly able to consciously read along. In 2001, sequencing efforts resulted in a first draft of human genome. Since then, the capacity of the DNA reading machines doubled on average every six months. While the first human genome sequencing project took years of worldwide collaboration, it is available as a service nowadays such as at the Genomics Core, where the equivalent of tens of genomes is sequenced every 10 days.
Each sequencing run gives rise to a few terabytes of raw data that, using bioinformatics techniques, must be processed in time, before the next bunch of data arrives.
I will discuss bioinformatics techniques commonly used in the Genomics Core and that have a chance to survive another generation of sequencing machines.
An important feature of these techniques is that they generate and distribute sub-tasks and thus maximize the use of computer clusters.
Dr. ir. Pieter Mestdagh
Center for Medical Genetics,
Ghent University
Integrative transcriptomics to study
non-coding RNA functions
Over the last years, non-coding RNAs (e.g. microRNAs and long non-coding RNAs) have emerged as an important layer of the transcriptome. In order to elucidate their function in disease biology, multiple tools have been developed, ranging from miRNA target prediction algorithms to the more advanced integrative genomics approaches. Through the combination of multiple layers of information, integrative genomics allows a more accurate and comprehensive assessment of non-coding RNA functions in human disease. In this presentation, I will discuss different approaches on how to combine multi-level transcriptome data in order to functionally characterize non-coding RNA networks.
Dr Yvan Saeys
Machine Learning and Data Mining
group, Bioinformatics and Systems Biology Division, VIB-UGent
Department of Plant Systems Biology
Large scale machine learning challenges for systems biology
Due to technological advances, the amount of biological data, and the pace at which it is generated
has increased dramatically during the past decade. To extract new knowledge from these ever
increasing data sets, automated techniques such as data mining and machine learning techniques
have become standard practice.
In this talk, I will give an overview of large scale machine learning challenges in bioinformatics
and systems biology, highlighting the importance of using scalable and robust techniques such as
ensemble learning methods implemented on large computing grids.
I will present some of our state-of-the-art tools to solve problems such as biomarker discovery, large scale network inference, and biomedical text mining at PubMed scale.
HIVlab, Department of Clinical Chemistry,
Microbiology and Immunology – UGent
Pathway analysis: example from the bench
To be completed
Prof. dr. Lennart Martens
UGent - Department of Biochemistry, Faculty of
Medicine and Health Sciences, VIB - Group Leader Computational Omics and Systems Biology Group (CompOmics),
Department of Medical Protein Research
High-throughput proteomics: from understanding data to predicting
them
In proteomics, as in any high-throughput omics field, the rate of data generation has increased dramatically, yielding very large datasets that require substantial processing to render them useful and interpretable. Key concepts here are data management, data-bound analysis algorithms, and user interface design. But we do not need to limit ourselves to only the interpretation of experimental results. By combining data from across many (unrelated) experiments, we can gain substantial knowledge about the strengths and limitations of our technological approaches. High-throughput methods however, rarely serve as the endpoint for research. As exquisite parallel hypothesis
testers, these approaches can quickly highlight promising follow-up targets for more detailed study. Yet moving from discovery to targeted analysis requires much more in-depth understanding of sample and methodology, which is where the insights gained from large-scale data analysis come into play. Armed with this knowledge, we can begin to predict experimental outcomes based on specific hypotheses, thus effectively creating tests or assays that can be used in focused validation experiments.

Registration is free but obligatory, by ….
Venue: Congress Center Het Pand
By public transport
- from train station Gent St-Pieters:
tram 1 (every 6 minutes), get off at Korenmarkt
- from train station Gent Zuid:
bus 16, 17, 18 of 19 (every 15 minutes), get off at Korenmarkt
By car
Follow parking route signs to parking P7 Sint-Michiels, less than 150 meters from Het Pand.