All posts by admin

Hap world map?

A new study of 12 Mb DNA sequence in 927 individuals representing 52 populations now finds good portability of of tag SNPs between the 4 hapmap groups and any of the 52 populations (except some African populations like the Mandenka, Bantu, Yoruba, Biaka Pygmy, Mbuti Pygmy and San). The paper has some exceptional well done graphics – and I am quite happy that the resolution of European nations leaves some gaps for our forthcoming ECRHS papers (a poster had already been on display at the 3rd Annual International HapMap Project in Cambridge, Massachusetts).

“Die Botschaft hör’ ich wohl, allein mir fehlt der Glaube” (Goethe, “I hear the message well…”). The usefulness of tagSNPs in disease association studies still remains to be shown (I still renember comments like cr.. map). At present I neither believe in rare variants nor in common common variants but a permanent reshuffling of rare, frequent and highly abundant variants. Yea, yea.

 

CC-BY-NC Science Surf , accessed 26.07.2026

Da steh’ ich nun, ich armer Tor!

[sorry in German, one of my favorite poems, Goethe – Faust you are right]

Habe nun, ach! Philosophie,
Juristerei und Medizin,
Und leider auch Theologie!
Durchaus studiert, mit heißem Bemühn.
Da steh’ ich nun, ich armer Tor!
Und bin so klug als wie zuvor;
Heiße Magister, heiße Doktor gar,
Und ziehe schon and ei zehn Jahr
Herauf, herab und quer und krumm
Meine Schüler an der Nase herum â€"
Und sehe, dass wir nichts wissen können!
Das will mir schier das Herz verbrennen.
Zwar bin ich gescheiter als alle die Laffen,
Doktoren, Magister, Schreiber und Pfaffen;
Mich plagen keine Skrupel noch Zweifel,
Fürchte mich weder vor Hölle noch Teufel â€"
Dafür ist mir auch alle Freud’ entrissen,
Bilde mir nicht ein, was Rechts zu wissen,
Bilde mir nicht ein, ich könnte was lehren,
Die Menschen zu bessern und zu bekehren.
Auch hab’ ich weder Gut noch Geld,
Noch Ehr’ und Herrlichkeit der Welt.
Es möchte kein Hund so länger leben!

 

CC-BY-NC Science Surf , accessed 26.07.2026

Calibrate!

This is a quick link to Eye of Science, a website with impressing micro photographs. Calibrate your monitor first at sriker, then goto Eye of Science. Oliver Meckes is quoting Albert Einstein

People should be ashamed to use the wonders of science and technology if they don’t know any more about it than a cow knows about the botany of the grass it relishes in eating.

 

CC-BY-NC Science Surf , accessed 26.07.2026

Men r’sponding to women

We know much about the differences between men and women – the X is the default pathway and the Y under the microscope looks as worn down and “misshapen as a stubbed-out cheroot“. There turns out to be something really new. So far all effects of Y genes on sex determination have been attributed to SRY, the testis determining gene (NR0B1, FOXL2 and WNT04 are probably ovary-determining).
The careful analysis of an Italian pedigree now described a new gene that can reversal XX to male when being disrupted: It is R-spondin 1 (or RSPO1), a growth factor that may act through ß-catenin stabilization and synergize with Wnt.
Do you know renember the nice cartoon of the Y chromosome with the HUH? selective hearing loss ;-) it is finally RSPO1. Yea, yea.

 

CC-BY-NC Science Surf , accessed 26.07.2026

Materials and Methods

Usually “materials and methods” section is in the second paragraph; some journals put it also at the end of a paper. As a reviewer I have always insisted that this heading should be extended to “Patients, materials and methods” while in epidemiology we frequently use the term subjects (BTW epidemiologists have a rather militaristic vocabulary: recruited, cohorts :-) ). The anonymous reviewer of a previous paper now pointed out: Use the phrase “participant” throughout and not “subjects” which has reductionist connotations. I promise to use “participants” from now on, yea, yea.

 

CC-BY-NC Science Surf , accessed 26.07.2026

Peer production

firstmonday has an interesting article about the limits of self-organization and “laws of quality”. Given 52 million tracks in the Gracenote database, 1 million entries in Wikipedia and 17,000 books in project Gutenberg, Paul Duguid throughly examines the two laws of quality

  • Linus law: “given enough eyeballs, all bugs are shallow” which means that almost every error will be discovered and ultimately fixed
  • Graham law: “people just produce whatever they want; the good stuff spreads, and the bad gets ignored”

Although more professionalized, similar principles operate in science. With these large genetic studies, I have the feeling that most errors occur at the interfaces, during hand-shaking of disciplines. There are certainly only a few people that can design a study, examine a patient, go to the laboratory, analyze and annotate the data and publish them. This means that even many eyeballs can not look around the corner and that it will take many years for the “good stuff to spread”. Yea, yea.

 

CC-BY-NC Science Surf , accessed 26.07.2026

Pharmacogenetic tests on the market

Certainly one of the best web resources for pharmacogenetics is the PharmGkB database that collects all kind of data about the relationships among drugs, diseases and genes. Of course you could sequence your genome or run expression profiling on a liver sample. However, you are probably here to find out what (serious!?) pharmacogenetic tests are already on the market.

Much can be said about the usefulness of such tests; I have doubts if there will ever be such personalized treatment as I can foresee some logistic problems to validate it ;-) More likely are group based therapies, maybe restricted to geographic ancestry. Here is a (first and very) preliminary collection of commercially available pharmacogenetic tests:

  • CYP2D6, CYP2C9 and CYP2C19 collectively account for about 40 percent of drug metabolism mediated by cytochrome P-450 (Roche). The AmpliChip CYP450 Test is the world’s first pharmacogenetic microarray-based test approved for clinical use. CYP2D6 metabolizes codeine into morphine. A variation in CYP2D6 varies with race and leads to a lower elimination rate of the antidepressants Prozac (a selective serotonin reuptake inhibitors); the alternatively used drug Celexa is metabolized by CYP2C19 (as well as omeprazole). Other examples include clopidogrel (metabolized by CYP3A4) and cyclophosphamide (by CYP2B6) and vitamin K (by CYP2C9)
  • NAT2*5A, NAT2*6A, NAT2*7A/B and NAT2*14A carriers are rapid and slow acetylators for example of isoniazid or procainamide (Roche)
  • HER2+ women may get herceptin (Roche, Bayer, PathVysion)
  • TP, DPD are the rate limiting catabolic enzymes of 5-fluorouracil metabolism (Roche)
  • Mitochondrial A155G variants are tested for aminogylcoside side effect (Humatrix)
  • A Warfarin sensitivity test will be in clinical use next year (Kimball Genetics). It will test for variations in CYP2C9 and VKORC1
  • An UGT1A1 gene variant is associated with leukopenia if prescribed camptosar, a drug for colon cancer (Oncoscreen)
  • A TPMT variant is associated with slow metabolism of 6-mercaptopurine, used in the treatment of childhood leukemia and inflammatory bowel diseases (Pharmaco-Gendia)
  • Epigenomics is currently developing tests based on DNA methylation
  • Tyrosine kinase inhibitor gleevec inhibits the ABL, ARG, SCF/KIT, and PDGFRA and PDGFRB kinases in CML. Mutations in ABL can arise as secondary mutations in previously sensitive leukemias (Pharmaco-Gendia)

Needless to say that I have excluded here specific HIV mutations that may induce resistance to particular drugs (as I learned last week on a bioinformatics meeting here by Thomas Lengauer). I have also excluded all kind of sex-specific marker (e.g. SRY testing) and the whole nutrigenomics stuff.

Who knows more, for example about lansoprazole effectiveness, UGT1A9 and mycophenolic acid, UGT1A1 and irinotecan, COMT genotype and amphetamine response, pharmacogenetics of COX-2 inhibitors, and GRP78 responsiveness to chemotherapy? Is there any commercial test available for these genes? It seems that somebody should start a wiki on that, yea, yea.

Addendum 31-12-09

Here is another gene list; only 6 tests have been approved by the FDA; Nature reports about Oncotype DX and Prostate Px as well as MammaPrint. See also an UK based paper

HLA-B*5701 was most commonly tested to identify those at risk of abacavir hypersensitivity among patients with HIV. A number of barriers to testing were identified, including lack of clinician knowledge and a lack of scientific evidence.

 

CC-BY-NC Science Surf , accessed 26.07.2026

INDELligent

I am detailing in a forthcoming paper in “Allergy”, that the contradicting results found with ADAM33 (the first positionally cloned asthma gene) probably results from a rather poor design of all follow-up studies.
It does not make so much sense to repeat over and over the same few SNP marker; instead a full resquencing of the linkage region would be necessary. From the analysis of public LD maps it is even possible that neighboring genes may be responsible for the observed associations.
I have also doubts if the SNP-centric view is always leading to success. BTW there is a new database of over 400,000 non-reduandt indels of which 280,000 are validated by comparison with other human or chimpanzee genomes (see Mills et al., the indels are available in dbSNP under the “Devine_lab” handle).

 

CC-BY-NC Science Surf , accessed 26.07.2026

Blog together

Just read the entry of Janet D. Stemwedel about the 2007 North Carolina Science Blogging Conference, a “free, open and public event for scientists, educators, students, journalists, bloggers and anyone interested in discussing science communication, education and literacy on the Web.” Do not miss her tag “Shameless self-promotion” ;-) Three editors from The Lancet will take up to 49 registrants. Who can give me a free ride from Munich (bonus miles welcome)?

 

CC-BY-NC Science Surf , accessed 26.07.2026

The journey is the reward

The analysis of a large dataset can be done in many different ways. At least in my experience documentation is getting confusing after many weeks of work – what has been done at which time? Can I reproduce my earlier findings? Where are the latest figures? What remains to be done? …

This is an ever increasing problem. Some information is documented in lab books, other in clinical journals, or even sitting on a server database where it may change over time. I have therefore described my own documentation procedure in a PDF paper that is online at this site.

With a quick view at the figure you will see that I am (mis-)using spreadsheets for documentation Please note (1) the drop down arrows at date and task that can be used for selecting lines even of long lists and (2) that tabs can used for fast switching between results (3) one tab contains all analysis scripts (4) and one tab all links.

spread.png

 

CC-BY-NC Science Surf , accessed 26.07.2026

Best of two worlds

Finally, linkage and association data can be used together after downloading new software using genotype inference.

It reduces the number of genotyping reactions and increases the power of genome-wide association studies. Our method combines sparse marker data from a linkage scan and high-resolution SNP genotypes for several individuals to infer genotypes for related individuals.

Sure, we

Geo IP identification

The Geo IP database is available at Maxmind and allows to trace your home city from IP addresses. Here is a quick and dirty script to upload the Geo IP data into MySQL:

|wj_geo.cmd|

I would also put an index on loc_id. Finally the database should be available as

SELECT city
FROM GeoLiteCity INNER JOIN GeoLiteCityBlocks ON GeoLiteCity.locID = GeoLiteCityBlocks.locID
WHERE $myIP >= startIpNum AND $myIP <= endIpNum;

where $myIP is calculated as

substr($_SERVER['REMOTE_ADDR'],0,3) * 16777216 +
substr($_SERVER['REMOTE_ADDR'],4,3) * 65536 +
substr($_SERVER['REMOTE_ADDR'],8,3) * 256 +
substr($_SERVER['REMOTE_ADDR'],12,3)

 

CC-BY-NC Science Surf , accessed 26.07.2026

A low-cost system for a PDF literature archiv I

Getting a scientific paper on your harddisk is quite simple. I am using a Fujitsu Scan Snap that can process a single page in a few seconds. The resulting PDF needs to be further tweaked by OCR recognition like ABBY FineReader (I couldn’t find any good open source alternative). FR will leave your PDF intact while adding recognized text as an overlay (or “underlay”). Unfortunately FR does not support batch processing but your OS will do by using a windows scripting engine like CLRscript. We also need a tool to extract a text file from the modified PDF. A good choice is pdftotext — look at the sourcecode and the DRM discussion before compiling it with a compiler like Cygwin. The following perl script doesn´t do anything than traversing your target directory and creating a batch file. As filenames offered by publishers are rather strange, I would first start to create some clean file names by replacing all spaces and brackets with something innocent like underscores.
perl.exe ocr.pl rename h:\pdf\2008\*.*
Now we create text files from the PDFs (usually done better by XPDF than directly by GDS).
perl.exe ocr.pl extract h:\pdf\2008\*.pdf
The resulting textfiles may be inspected: very small file sizes usually indicate no valid extraction and should be deleted before starting the OCR step as OCR is only done when text files are missing.
perl.exe ocr.pl ocr h:\pdf\2008\*.pdf
In the last step you may want to repeat the extract step.

ocr.zip
|wj_ocr.txt|

 

CC-BY-NC Science Surf , accessed 26.07.2026