Showing posts sorted by relevance for query emboss. Sort by date Show all posts
Showing posts sorted by relevance for query emboss. Sort by date Show all posts

Sunday, 27 January 2013

The Horseburger Protocol

But let's leave the pastures of the past and - revenons nous a nos moutons - get back to the day job.  I wrote earlier about my molecular biology lab section.  We're spending the next tuthree weeks using the enzymes (restriction endonucleases -REs) which cut DNA at particular sequences of bases pairs [e.g. an enzyme from our homely commensal E.coli, called EcoRI, cuts DNA at GAATTC].  We have a CSI protocol where DNA from 4 suspects and a crime scene sample is treated with a couple of REs, and the fragments run out on a gel. The pattern generated by The Perp will then be compared to that from the butler, the husband, the lover and the nephew-who-will-inherit-everything.  In circulating this proposal, the Course Director put a footnote "and maybe Bob would have some bioinformatics ideas on this".
As I've spent most of the last 25 years doing 'the storage retrieval and analysis of biological sequence data' aka bioinformatics, you'd think I'd have loads to contribute.  Well I didn't, and frankly I was a teeny bit teed off that They would expect me to do anything beyond float-not-drown in the deluge of my new responsibilities. But two days later, I was driving to work through a literal deluge and had the wireless off, so I could concentrate on the traffic, and . . . I had an idea <ching!>.
There has been a scandal recently in Ireland over the revelation that cheap frozen '100% beef' burgers have been packed out with horse-meat.  There is 'no health hazard whatsoever' in this but if the producers finagle the label, what else - that we cannot detect - might they finagle?  Then again, as the rhetoric has it: what do you expect if you spend €1.49 on a dozen burgers??
ANNyway, in my last job as a Comparative Immunologist, I spent a lorra time comparing DNA sequences from different species - a lot of them mammals.  So my <ching!> was for the students to e-go to NCBI http://url.ie/guzz in Bethesda MD and fetch out a pair homologous sequences: one from Equus caballus and the other from Bos taurus. They can then search the sequences of the genes for GAATTC and the other RE sites.  A random pair of homologous genes from two species of different mammalian Orders might be 15% different, so there should be differences in the restriction pattern, even if you have to try a couple of different enzymes.  Appropriate-technology-me likes this a lot because you can do the whole analysis with a highlighter.  But you can't expect students to carry out tedious repetitive tasks nowadays - we have software for doing that.  So I'm suggesting that they use Fuzznuc http://url.ie/gv01 from the EMBOSS suite of binfo software, to do the grunt work for them. The Horseburger Protocol - sounds like a racey Robert Ludlum thriller.
Haven't a clue what gene to choose?  Here's a starter list: ANXA1, BRCA1, CFTR, DRD1, EFNA1, FZD3, GSTK1, HSPA1A, IFNG, JBS, KLHL3, LPAR2, MTNR1A, NECAB2, OPRD1, PELI3, QARS, RNASE1, SOCS3, TSHZ1, UBXN8, VCF, WAS, XDH, YEATS2, ZBP1.  Don't like any of those?  There are 20,000-26 = 19,974 others to choose from. Go to!

Wednesday, 9 December 2020

Small country Big data

 As my ambition genes were shot off in the war, I was never very successful in science. If I'd cared more about getting ahead, I'd have been better at the hard work of publishing papers and might thereby have gotten tenure at Harvard Little Rock. I applied there immediately after my PhD - not short-listed! But I did have a lurching mostly on-track career in science for 50 years. I was much better at infrastructure than the cutting edge: filling in the pieces behind pioneers to firm up the foundations for the next Great Leap Forward.  In the 1990s I was the Director-and-Sole-Employee of The Irish National Centre for BioInformatics INCBI which made me the Irish node for EMBnet, an EU-funded quango of bioinformatic support centres. The vision for such infrastructural support of Ireland's bioscience revolution was there but the money wasn't. As the cash dribbled away with the century, I was partway along a cunning plan to get a suit and a carry-on bag and sell seats for eBioinformatics, a commercial platform of getting access to the software and databases for cracking the code of the human and other genomes. Tim Littlejohn, the Australian founder and CSO, was preparing to appoint me sales-and-support Europe. It never happened: eBioinformatics didn't grow fast enough for sufficient Venture Capital to be interested and a handful of  Tim-and-my colleagues in EMBnet developed a clunkier but sufficiently effective free platform called EMBOSS which took the wind out of the sails of any commercial venture.

At about the same time another wannabee Merchant Venturer called Kári Stefánsson jumped ship from actual Tenure at Harvard and returned to Iceland [L þinnvellir rift valley] to investigate the genetic basis of MS. He founded a company called deCODE with a business model to sequence the DNA of everyone in Iceland and cross-reference the genetic variants against their disease status. That way, deCODE might be able to identify the key genes for MS predisposition and start to develop effective therapies against the degenerative disease. And while you're about it why not cover other line-items on the medical charts: cardiovascular, diabetes, stroke, cancer . . . I had a brief fantasy about leaving Ireland and whoring myself out in Reykjavik but nothing came of that . . . because ambition genes and because we'd just bought a fantasy farm and the fairies had delivered two infant girls to look after the hens.

deCODE was in the Covid News last week because, for a small country, Iceland has delivered Big Data. There are 50% more people in County Cork than in Iceland. At the beginning of the pandemic Stefánsson offered the facilities at deCODE to the Icelandic Directorate of Health after the equipment started to melt-down at te labs in Landspitali — The National University Hospital of Iceland. deCODE, as a big science genomic centre, has a lot of capacity and the DoH picked up the tab for anyone with the least hint of covid to get tested and resulted within 30 hours. So about half the population got tested; those who were positive were monitored by public health so we have a bunch of outcomes that should be (with caution) translatable to wherever you live.

  • 43 % of people who are infected with the virus are without symptoms
    • that is exactly the same rate as found in the N. Italian city of Vo which locked down and track-and-traced in the Spring
  • the most common symptoms are
    • muscle ache
    • head ache
    • unproductive cough
    • not rise in temperature
  • The fatality/case rate is 0.3% !
    • seasonal 'flu is 0.1%
    • median age for Iceland 37yrs is low, so fatality will be higher in countries with more octogenarians.
  • Population death rate 7 /100,000 less than 10% of that in USA or Britain
    • Ireland: 40/100,000
    • Norn Iron: 50/100,000
Tourism supports the economy in Iceland and tourists are being allowed back to visit the place where the Atlantic is spreading its plates [whc prev]. After an outbreak blip traced to two scofflaw tourists, stricter control measures are now required: a) either self-quarantine for 14 days on arrival or b) do two screening tests: one on arrival + one  five days later. 20% of people negative in the first round tested positive on second go! These tourists are pretty hot with the Covid. Which only says that the prevalence is high among the demographic which travels.

Wednesday, 14 May 2014

It's in code

Whoop whoop - nerd alert.
When I wrote about the Great Western Binfo meeting I attended last November, I prefaced the piece with a too-clever-by-'arf
UUUAGAAUCGACGCAUAC AUCAAC ACCCACGAA UGGGAAUCCACC
thinking it would serve as un p'tit amuse-cerveau for binfoes before the meat of the article. I suspect that it served, to the nearest whole number, solely as un p'tit amuse-Bob.  Last week I had the honour to host the 2014 version of the meeting and I put
GUNAUHRAYGAR
in the header of the programme by way of continuing the tradition.  It came up in conversation during the afternoon coffee break when I was chatting with two physicist-turned-binfoes.  Somebody who was earwigging turned round and said "Oh I assumed that was a typo".  Well, really!  I know that The Blob is sprinkled with errurs of spelinge and I often blush when I re-read my e-mails, but if I write something that sounds like one of two trolls 'wrestling' then I do so deliberately.  Because the p-t-bs are both fightin' sharp and quick on the uptake, it didn't take them long to crack the code.  That led nicely into a discussion of the IUPAC codes for DNA/RNA bases and the amino acids that go to make proteins.

Cypherists are quite frustrated by the fact that, while there are 26 letters in the English alphabet, there are only 20 amino acids that commonly appear in protein sequence
A AlaC CysD AspE GluF PheG GlyH HisI IleK LysL Leu
M MetN AsnP ProQ GlnR ArgS SerT ThrV ValW TrpY Tyr
BJOUXZ are all missing which is a drag because the list includes half the vowels. So you can find ELVIS in the protein database but you can't find BOB. The AAs are translated from RNA read in triplets, so the 12 letters at the top could be a genetic code to represent a four-letter word. Except that it has letters in addition to the regular ATCG (DNA) or AUCG (RNA).  What am dem and why dey needum?  They are usually referred to as ambiguity codes.  The "four" bases are of two chemical types: puRines (lArGe) and pYrimidines (CUTe or small).  Sometimes you don't know, or cannot specify, which purine is present so you write R.  Likewise the bases pair with each other using either 3 hydrogen bonds CG or 2 A=T.  The former is stronger than the latter so C or G is denoted S trong and A or T as W eak. N on the other hand means any base.  You rarely need the others but the finite number of ambiguous possibilities each has its one-letter IUPAC representation: Isoleucine - Ile - I is translated from any codon that starts with AU and doesn't end in G.  IUPAC-speak for this is H (not G).
AUUIle
AUCIle
AUAIle
AUGMet
Amino acids also have their ambiguity codes and for a good chemical reason. When you wish to characterise a protein one technique is to break it down into its component amino acids and count their relative abundance - this can be surprisingly informative about which protein you have.  But when you hydrolyse the peptide bond between each pair of AAs, you also hydrolyse the chemically very similar amide bond that differentiates the amino acids aspartic acid Asp from asparagine Asn and glutamic acid Glu from glutamine Gln. A complete hydrolysis with 6N HCl thus results in only 18 different components and the ambiguity is written thus Asp or Asn = B and Glu or Gln = Z.  Just as N represents any base in DNA, so X represents any AA in protein.


That's handy for geeky coders because it brings two more letters into the possible alphabet for writing English words in codons. In December I was well impressed by a couple of the Masters of Immunology knowing that there was a 21st amino acid selenocysteine which is incorporated into proteins in special circumstances in particular species by subverting UGA, one of the stop codons. Selenocysteine is interesting because it looks exactly like cysteine except that a selenium atom replaces cysteine's sulphur. And that's interesting because oxygen, sulphur and selenium are all in the same column of the periodic table of elements - ie. have similar chemical properties. A couple of days later one of the MScs e-mailed me to say "and let's not forget pyrrolysine".  That's another oddity that was discovered in 2002 in the methyltransferase gene of Methanosarcina barkeri.  It's since been found in methyltransferases from other species where it forms an essential part of the active site.  Like selenocysteine, it is coded by a subverted stop codon, in pyrrolysine's case UAG.  This is handy because IUPAC have elected to represent selenocysteine as Sec or U and pyrrolysine as Pyl or O, so we can write pretty much anything we want in codon-speak because the only letters uncodable are minority interest J and X.  But even those are sorted because X (any AA) could be represented by codon NNN and you could use I for J like the Romans.

Emboss backtranseq could help you if you want to send a secret message to your geeky mol.bol girlfriend like

ATC TTCGCCAACTGCTAC TACTAGTGAAGG GACAGGGCCATCAACAGC 
AGCATGGCCAGGACCTAC CCCGCCAACACCAGC
and ExPaSy will help her translate the DNA back into English.  I hope you know what to do when you do get together - and no, there's more exciting things to do with her than "pull an all-nighter playing Dungeons & Dragons".