All posts by admin

All roads to NLM

This is not just an addendum to my previous post free-for-all or to number-cruncher: the 12 Dec NIH press release links to a new and exciting database

NIH Launches dbGaP, a Database of Genome Wide Association Studies
The National Library of Medicine (NLM), part of the National Institutes of Health (NIH), announces the introduction of dbGaP, a new database designed to archive and distribute data from genome wide association (GWA) studies. GWA studies explore the association between specific genes (genotype information) and observable traits, such as blood pressure and weight, or the presence or absence of a disease or condition (phenotype information).

Addendum

29-5-07 dbGaP suffers from some broken links but content improves!

dbgap.png

 

CC-BY-NC Science Surf , accessed 27.07.2026

Christmas present – your digital book copy

Maybe you don´t want to wait until Google Scholar has it; maybe you are interested in a higher quality: Here is the web address digiwubu.gdz-cms.de at the Niedersächsische Staats- und Universitätsbibliothek Göttingen where I once studied theology.
There should be no problem to scan any book published before 1900, however you can ask also for books published later than that date. Costs will be are 0.25 € per page plus 5 € for handling and shipping a CD.
“Google Books Library Project” is currently scanning in Harvard, Stanford and Oxford >15 million volumes – German scan factories are in Göttingen and Munich. Yea, yea.

 

CC-BY-NC Science Surf , accessed 27.07.2026

Exodus of science from Germany after 1933

A book that crossed my desk only very recently is about the exodus of science from Berlin after 1933. As a child I never understood the second commandment when God said to Mose that

Ex 20:4 I am a jealous God, punishing the children for the sin of the fathers to the third and fourth generation of those who hate me, but showing love to a thousand {generations} of those who love me and keep my commandments.

I always thought the idea to be unfair to be under collective guilt. Nevertheless when reading this book (published already in 1994 by Walter de Gruyter) we get a deeper meaning how science is affected for many generations by the displacement of the most prominent scientists.
Particular important in this book is the first chapter of Hubenstorf and Walther that highlights the situation in Berlin. Following world war I, Berlin had become the indisputable center of most scientific disciplines in the German speaking territory. In some disciplines Leipzig, Vienna or Munich may have been competitors, economics had been strong in Kiel, mathematics and physics in Göttingen, history in Marburg, however, for most scientists Berlin had been the highly desired “endpoint” of their career. he Friedrich-Wilhelms university had been the largest university, but there have been many more science organizations like Technical University Charlottenburg, Deutsche Hochschule für Politik, Preußische Akademie der Wissenschaften and Kaiser-Wilhelm-Gesellschaft (that is covered in more detail in another excellent chapter).
Medicine has been hit hardest during the Nazi period by having a large number of Jewish scientists. The resulting repercussions in the realm of science are described at different levels. Starting with a typology of transformation of scientific institutions, the establishment of new disciplines and the establishment of a military science sector, the autors give many historical details about the ways of scientific publishing or the organization of displaced scientists.
The cancer research department of the medical faculty fired 12 or 13 scientists; the hygiene institute dismissed 8 of 12 scientists including later Noble prize winner Erwin Chargaff. Hospital Lankwitz fired all physicians, Neukölln 67%, Freidrichshain 62% and Moabit 56%.
It is a terrible story – you can read how the editor of the Deutsche Medizinische Wochenschrift Paul Osswald Wolff was replaced by a Nazi supporter. Karger publisher even moved from Berlin to Basel (where they still reside today).
As any good science is strongly connected to teaching it may be understood that breaking this tradition has lead to a punishing of the children for the sin of the fathers to the third a fourth generation. Science politicians may even recognize the downside of spending money into science: they will be blessed by thousands of generations.

p1000238.JPG

 

CC-BY-NC Science Surf , accessed 27.07.2026

R parallel computing

Following several unsuccessful attempts to implement a parallel computing platform for R statistical software, I am showing here my current approach that is largely influenced by a recent paper on cluster programming in c’t 6/06 by Oliver Lau (sorry, no online version). My primary interest is with the R library snow (or snow-ft) that offers the function clusterApplyLB. This function is all I need for my R programs.
Now it gets more complicated: library(snow) depends on library(Rmpi): Hao Yu has an excellent description at www.stats.uwo.ca/faculty/yu/Rmpi how to set up the mpi layer with MPICH2. I am currently experimenting with DeinoMPI a closely related high performance Windows interface. According to its developer David Ashton it has the following advantages

First, DeinoMPI does not require MPI applications to be started by mpiexec in order to call MPI_Comm_spawn so you could load Rmpi from the Rgui.exe without having to bother with calling mpiexec. Second, DeinoMPI loads the user profile when starting applications so if you query the user’s temporary directory you will get the user specific path and not the Windows system temp directory. Third, DeinoMPI handles arguments with spaces correctly if you quote them so you can pass environment variables with spaces in them. Fourth, DeinoMPI allows you to use the MPI Info object to pass extra options to MPI_Comm_spawn like drive mappings. So you could create an MPI_Info object and set wdir=z:\ and map=z:\\server\share. Then pass this info object in with the MPI_Comm_spawn command and you could map a network drive and launch an executable from this drive.

So far the Rmpi package is compiled for MPICH2 (not DeinoMPI) so it won’t run with only DeinoMPI installed but there is a good chance that this will change in the near future.
Further useful references are in the R newsletter 2003, p21 cran.r-project.org/doc/Rnewsand a paper in the UW Biostatistics Working Paper Series on “Simple Parallel Statistical Computing in R” by Anthony Rossini and LukeTierney.
BTW, haplotypes of the hapmap project were computed on a 110 node cluster provided by both Peter Donnelly’s Mathematical Genetics Group www.stats.ox.ac.uk based at the Oxford Centre for Gene Function and by a 128 node compute cluster provided by the Oxford e-Science Centre e-science.ox.ac.uk as part of the National Grid Service[to be cont’d…].

mpi1.png

 

CC-BY-NC Science Surf , accessed 27.07.2026

3D LD

While waiting for genomewide SNP data to be re-partioned into LD blocks I found this page with some neat progamming tricks. It is part of the dissertation of Ben Fry / MIT about computational information design. Page 74 ff has a history of redesigning the widely used haploview pogram.

The design of these diagrams was first developed manually to work out
the details, but in interest of seeing them implemented, it was clear that
HaploView needed to be modified directly in order to demonstrate the
improvements in practice. Images of the redesigned version are seen
on this page and the page following. The redesigned version was even-
tually used as the base for a subsequence ‘version 2.0’ of the program,
which has since been released to the public and is distributed as one of
the analysis tools for the HapMap [www.hapmap.org] project.

3d-ld.png

 

CC-BY-NC Science Surf , accessed 27.07.2026

How can we know?

A recent paper in Nature reported

Tissue samples were obtained from one of the following sources: Asterand, Pathlore, Tissue Transformation Technologies, Northwest Andrology, National Disease Research Interchange and Biocat. Only anonymized samples were used, and ethical approval was obtained for the study from Ärztekammer Berlin and the Cambridge Local Research Ethics Committee. […] Human primary cells were obtained from Cascade Biologics, Cell Applications, Analytical Biological Services, Cambrex Bio Science and the Deutsches Institut für Zell- und Gewebeersatz.

How did “Ärztekammer Berlin” or “Cambridge LREC” evaluate ethical performance of these companies? Or did anonymity automatically guarantee ethical research? Or is it just a formal requirement to mention ethics? Or …?

 

CC-BY-NC Science Surf , accessed 27.07.2026

Open culture podcasts

As a frequent traveller I like podcasts. Here is a quick link to Open culture that have a huge university podcast collection including many foreign language selections (Boston College, Bowdoin College, Collège de France, Duke University Law School, Harvard University, Haverford College – Classic Texts, Johns Hopkins, Northwestern University, Ohio State, Princeton University, Stanford University, Swathmore College, University of California (the best collection), The University of Chicago, The University of Glasgow, The University of Pennsylvania, The University of Virginia, The University of Wisconsin-Madison, Vanderbilt University, Yale University and Ecole normale supérieure). If you don´t like proprietary formats you need to find the good and the bad apples.

 

CC-BY-NC Science Surf , accessed 27.07.2026

Evolution in fast motion

Nature genetics as an advance online publication about comparative genome sequencing of E. coli where 13 de novo mutations in 5 strains were monitored over 44 d (or ~660 generations). It is a great study – not only because the author list includes one of my previous coauthors – but for giving a first insight about development of a mutation and fixing its allele frequency. Unfortunately, there is no flowchart and the methods are somewhat vague, what has been sequenced (or resequenced) in which strain at what time . In other words who are the winners? Did they manage that by their own strength or with a little help of some friends? Why rises the allele frequency always to 100% and what about some discrepancy of allele frequency and fitness? We will hopefully see more of these studies, yea, yea.

 

CC-BY-NC Science Surf , accessed 27.07.2026

Dr. med. Sigmund Rascher, KL Dachau

On my way to work I am crossing every morning in Dachau East the former Nazi concentration camp/Konzentrationslager (KL). Its a monument of inhumanity and the deepest point in the history of “science”. A large number of prisoners were abused by SS doctors for medical experiments; an unknown number of prisoners suffered agonizing deaths in the course of atmospheric pressure, hypothermia, malaria and other experiments.

photo by 11 Dec 06
p1000219.JPG

Having a longstanding interest in history (and even published on the 50th anniversary of the Nuremberg trials) I have now been very interested in a new book by Sigfried Bär, one of the outstanding German science writers “Der Untergang des Hauses Rascher”, a history of the life of Dr. Sigmund Rascher, anthroposophic scholar, medical student, DFG-scholar, minion of of Heinrich Himmlers, air pressure and hypothermia researcher at KL Dachau and finally prisoner who died by being shot in the neck.

Dr. Bär spent several years researching the life of this mass murderer. He contacted relatives of Rascher, looked at family photos, talked to people who knew Rascher and went to archives. This is a unique document showing the avidity of a researcher for recognition by scientific colleagues. Other books from my own library that I recommend:

pb300004.JPG

 

CC-BY-NC Science Surf , accessed 27.07.2026

Pathway to nowhere

I love pathway diagrams since I mounted the famous Biochemical Pathways of Boehringer in my bachelor flat. As far as complex disease genetics is concerned with many disease genes, an integration into a pathway context becomes critical. There are many attempts to extract this information from the literature and many companies that offer highly curated information (Biomax, Ariadne Genomics, Genomatix to name a few). Academia relies mainly on KEGG, the Kyoto encyclopedia of genes and genomes or Biocarta. Last week another pathway server appeared that is curated by the NCI and Nature magazine. Let’s have a look – I am currently working on Affymetrix 500K SNP annotation, yea, yea.

 

CC-BY-NC Science Surf , accessed 27.07.2026

How to detect your own CNVs

How to detect copy number variation (CNV) in your own genotype chip data, can be found in a companion paper of the recent Nature publication.
In the previous Nature paper the authors explained their algorithm to be based on k-means and PAM (partitioning around medoid) clustering, but it seems quite different. They call genotypes with DM (which seems to be already obsolete by the BRLMM, see a comparison at Broad and the AFFX whitepaper), then adjust heterocygote ratios by Gaussian mixture clustering, normalize and reduce noise before! merging NspI and StyI arrays. The software is at Genome Science, Tokyo. Yea, yea.

 

CC-BY-NC Science Surf , accessed 27.07.2026

On mice and men with asthma

A new series of pro- and con editorials in the Am J Res Crit Care Med discusses the question why in some instances mouse models have “misdirected resources and thinking”. You may have noticed that I have only rarely used animals for research; the authors of this editorial have collected empirical data on the exploding use of murine models. Despite their attractiveness from a technological point, they are often useless because

  • mice do not have asthma as even the most hyperresponsive strain does not show spontaneous symptoms
  • mice do not have allergy – although sensitization can be manipulated by high intraperitoneal allergen/adjuvant injection, this does not involve immediate and late airway obstruction.
  • immune reaction in mice is quite different – the interfering of some substances like vitamin D cannot be reliable tested, there is no pure Th1 and Th2 reaction in human and less stronger IL-13 response
  • mice typically can not be challenged with the complex (and interacting) human exposure – oxidant stress, viral infection, obesity, diet, smoke, pollutants, ….
  • time course is difficult to mimicking in the mouse, there is no longterm model
  • structure of mouse airways is different – there are fewer airway generations, much less hypertrophy of smooth muscle
  • inflammation in mouse is parenchymal rather than restricted
  • humans are outbred, mice are inbred
  • early microbial environment is different
  • many promising interventions of mice pathways failed in humans (VLA-4, IL4, IL5, bradykinine, PAF,…)

I am sure there are even more arguments – I suggest that the authors deserve the Felix-Wankel price.

Addendum

15 Dec 2006: The BMJ has 6 more examples about the discordance between animal and human studies: steroids in acute head injury, antifibrinolytics in haemorrhage, thrombolysis or tirilazad treatment in acure ischaemic stroke, antenatal steroids to prevent RDS and biphosphonate to treat osteoporosis.

19 Dec 2006: Another pitfall paper

31 Dec 2006: A blog on animal welfare

25 Apr 2007: Call for better mouse models

15 Jul 2018: Of Mice, Dirty Mice, and Men: Using Mice to Understand Human Immunology

16 Oct 2019 Allergy protection on farms – why also studies in mice could have failed

18 Feb 2023 The case for free running lab mice, see
Cell and Science paper

7 Jan 2023: A review concluding that The vitamin D system in humans and mice: Similar but not the same

12 March 2024: Limitations of the mouse model

  • Lack of clinical relevance - The OVA model often fails to accurately replicate human asthma and allergic diseases. OVA is not a common human allergen, unlike house dust mites, pollen, and pet dander.
  • Artificial sensitization - The model requires sensitization with alum, a strong Th2-skewing adjuvant, which does not reflect natural allergen exposure in humans.
  • Simplistic immune response - OVA exposure leads to a predominantly Th2 immune response, whereas human asthma involves a complex interplay of Th1, Th2, Th17, and innate immune responses.
  • Limited chronicity - Most OVA models induce short-term allergic inflammation but fail to reproduce the chronic airway remodeling, fibrosis, and structural changes seen in human asthma.
  • Non-physiological exposure - OVA is often administered at high doses via intraperitoneal injection or repeated aerosol exposure, which does not mimic real-life human allergen exposure.
  • Strain-specific bias - Different mouse strains respond variably to OVA, making it difficult to generalize findings.
  • Lack of environmental and genetic factors - The model does not account for genetic predisposition, pollution, infections, or microbiome variations that influence human allergic disease development.
  • Poor translational value - Many drugs that show promise in OVA models fail in human clinical trials, highlighting the model's limited predictive power.

 

CC-BY-NC Science Surf , accessed 27.07.2026

Why dog and cat can’t marry

or in more scientific terms: Why are F1 hybrids so often sterile or lethal?

The Dobzhansky-Muller theory says that there is an incompatibility between genes with reduced fitness that have diverged between species. So far nobody has ever observed a D-M gene but a new Science paper describes two genes that separate D. simulans and D. melanogaster: lhr (lethal hybrid rescue) and hmr (hybrid male rescue).

Well, dogs have 78 chromosomes, arranged in 39 pairs, while cats have 38 chromosomes, in 19 pairs – so no viable embryo can be produced.

 

Yea, Yea.

 

CC-BY-NC Science Surf , accessed 27.07.2026

For the first time two human genomes compared

Another “first discovery” in this nature genetics preprint although the analysis could have already been done some years earlier. The CNV specialists from Toronto now compare the Human Genome Project sequence with the Celera sequence – the gap between the two compilations was obviously bigger than the intra-sequence gaps. Of course both sequences are still mosaics from several individuals but the analysis nicely exemplifies how difficult it will be to compare the genome of two different human beings.
The authors employ a whole battery of alignment tools BLAT, MEGABLAST, GCA and A2Amapper. Of course results depend on the strategy, definition and implementation. As show by FISH analysis most of the discrepancies are true and can be classified into a few categories – insertions or deletions if seen from the second genome (has somebody ever thought about a minimal human genome?), mismatches and inversions. We are getting here a preview of the diagnostic workup in a patient in 2026. This blog contains forward looking statements while the responsibility rests solely with the reader. Yea, yea.

 

CC-BY-NC Science Surf , accessed 27.07.2026