San Carlos, California, United States
In my current role, I provide scientific, technical and strategic leadership on bioinformatics and genomics at Solvuu. I drive product and business requirements towards building Solvuu's agile data science platform for life sciences. Our goal is to take molecular labs into the future by providing the digital infrastructure for precision medicine. My previous responsibilities involved developing and implementing bioinformatics solutions to reduce dimensions in genome data, next-generation sequencing data, mass-spec proteomics data and microarray data. I was also responsible for ensuring bioinformatics data analysis in a GLP framework for products in regulatory science phase of their development. I have biological expertise and research interest in the areas of stem cell biology/regenerative medicine, cancer biology, clinical diagnostics, agricultural solutions, microbiome research, and plant biology.
Build bioinformatics tools for clinical diagnostics. Provide in silico reference materials (ISRMs) for NGS-based clinical assay validation, evaluate and validate the accuracy of variant calling methods to improve test accuracy.
Led the bioinformatics team on delivering consulting solutions for biotech and pharma customers. Oversaw the execution of both short-term and long-term NGS-based projects, ranging from routine analyses to highly complex discovery efforts in oncology therapeutics (target identification and validation, cell therapy), microbiome therapeutics, and metabolic & endocrine disorders.
- Provided scientific leadership on developing and implementing bioinformatics methods, data analysis/integration approaches towards Trait research. - Served as bioinformatics unit representative coordinating all bioinformatics and systems biology tasks (project requirements specification, experimental design, data analysis, interpretation and outcome dissemination) on two multi- year multi-function, multi-site collaborative projects. - Analyze customer/project requirements and developed focused, timely and efficient bioinformatics solutions to reduce dimensions from millions of data points to a manageable and biologically meaningful set of data points and actionable items towards Trait research. - Collaborate with a wide variety of BASF functions, and internal customers from Crop Protection Research, Plant Science, White Biotechnology, and Experimental Toxicology & Ecology towards understanding and solving their bioinformatics, systems biology and statistical data analysis needs. - Worked with external collaborators on evaluating and developing relevant tools for in-house projects and needs. - Implemented open source bioinformatics workflows to analyze marker gene data (16S and ITS) from NGS meta-profiling experiments and downstream statistical analysis towards Plant and Human microbiome characterization. - Projects contributed: Plant and Human microbiome characterization, Proteogenomics Project, Herbicide Tolerant Crops, Plant Yield and Stress, and Fungal Resistance Corn & Soybean.
- Developed and implemented bioinformatics pipelines and analyzed next-generation sequencing (NGS) data in deep sampling of transcriptomes (RNA-seq) towards data QA, expression-level estimation (gene/isoform-level), differential expression assessment (gene/isoform-level), sample QA, visualization for reference transcriptomes. Analyzed 5000+ RNA-seq datasets, demonstrated scale-up of pipeline to handle large datasets. - Developed computational pipeline to process NGS data for de novo characterization of transcriptomes without a reference, assess assembly quality and downstream analysis (transcript quantification, sample QC, differential transcript usage, identify sequence variants, coding region identification, functional annotation and enrichment analysis). - Developed computational workflows and led computational analysis efforts to integrate genome sequence data, NGS transcriptomics data and LC/MS-MS proteomics data using commercial and open source Unix-based tools towards new gene discovery and accurate gene structure annotation. - Compiled research reports on completed projects and communicated findings to customers through regular presentations and method demonstrations. - Worked on several time-bound projects from diverse functions and projects with a wide array of bioinformatics needs.
Research Interests: Integrating high-throughput data on the transcriptome, epigenome, methylome to understand how hematopoietic stem cells (HSCs) are generated during embryogenesis, and how their self-renewal ability and multi-potency are established through microenvironmental cues and nuclear regulatory mechanisms. - Worked with experimental bench biologists including graduate students, post-doc associates, staff and PIs on translating genome-wide, high-throughput data (NGS and microarray) into biologically meaningful results towards validation and testable hypothesis. - Implemented bioinformatics pipelines and provided data analysis support towards making sense of microarray expression data, RNA-Seq data and ChIP-Seq data. - Provided hands-on training in analyzing genome-wide expression data to non-specialists. Developed test datasets and tutorials.
- Provided high quality genome annotation through computational and manual analysis of structure and function of genes and other sequenced objects in the genome of model plant species Arabidopsis thaliana. This involved analyzing and interpreting sequence data, combining evidence from mapped sequence objects, comparative genomics techniques, examining experimental data and compiling results; assisting in developing improved formats and methods for community access to TAIR genome releases; soliciting community feedback and incorporating it into future releases. - Worked with software developers to maintain and improve pipelines for mapping a variety of sequenced objects (cDNAs, ESTs, high-throughput sequencing data, proteomics data, polymorphisms, markers, methylation data, microarray data, etc) to the genome. - Handled user requests on data analysis from plant researchers worldwide, presenting and publishing data. - Worked with the technical team to assist with the design of new or improved web interfaces and tools. - Helped conduct workshops, training sessions on genomic data analysis and using bioinformatics tools.
Research Interests: Genome biology: One of my research goals in Prof. Mark Gerstein's Lab was to identify and elucidate the role of various transcribed and translated sequence elements in eukaryotic genomes using computational approaches. Protein function: I was also interested in exploring how the physical, chemical and historical properties of proteins shape their structure, function and evolution. Projects: - Implemented bioinformatics methods for data analysis of whole-genome tiling microarray experiment and high- throughput sequencing data towards functional genomics of Rice and Arabidopsis. - Worked on and collaborated with other scientists on the modENCODE Project towards global identification of transcribed regions of C.elegans genome.