Tuesday, 12 June 2012

Of grants and grunts

I'm about to finish writing up the proposal of a research grant I'm applying for. It's about the use of Regression Discontinuity Design to evaluate interventions in primary care. 


The idea is quite clever, I believe: say that there is a clear rule to guide the way in which patients are (or aren't) given a treatment, and this rule is associated with a continuous variable. For example, say that you are given drug $d$ if your blood pressure is above some level $x$, but you are not if it is below. If this is reasonable, subjects just on either side of the threshold can, to some degree of approximation, be considered as randomly allocated to the treatment, thus mimicking controlled conditions.


Of course, this doesn't necessarily eliminate bias and confounding, but it goes quite some ways in limiting the impact of these problems. Also, it turns out that in clinical practice there are quite a few examples of situations that fit this framework. So we're proposing all sorts of cool investigations $-$ hopefully they'll like it and give us all the money we've asked, which I'll then use to buy new players for newly promoted Sampdoria.


That's for the grants part. 


Grunts are about: 

  1. the fact that I'm really getting fed up with reviewing the proposal and I'm really looking forward to the moment I'll submit it; 
  2. the fact that the waitress misplaced our pizza order, which means it took her forever to bring our food;
  3. English summer rain (which at the moment seems to have stopped, but the guy on BBC has just said there's more to come).

Friday, 8 June 2012

Euro 2012 predictions

In the run up to major football events, the guys at the Norwegian Computing Centre always prepare their predictions for the final results. They use a relatively simple model which predicts the number of goals scored by team $t_1$ when playing team $t_2$ as a function of a "baseline" scoring intensity, corrected for the relative difference in strength between the two teams.


This is not too dissimilar from the method that Marta and I used a while ago (by the way: interesting story. I was trying to have an excuse to bring more football into the household $-$ she got interested in the model, but no actual increase in the time spent watching games has ever been recorded). 


Our main point was that there is indeed correlation between the number of goals scored by the two competing teams, but instead of using complex forms for the likelihood (eg bivariate Poisson models), it can all be accounted for by extending the structure to a hierarchical model, based on independent Poissons that are connected through common structured effects. We did reasonably well:




(the black line is the observed dynamics in terms of points through the 2007-2008 season of the Italian Serie A, while the blue line is our prediction; for most teams the prediction is really good).


In our case, we were estimating results from national leagues, which I think is a bit easier, given that there are more data and that the seasons are longer, meaning that the "true" values tend to come up with a stronger signal (and stronger teams will be predicted to score more goals and thus win more games). There was an interesting issue with overshrinkage (check the model out).


In the case of international tournaments, prediction is a bit more complex; the tournament is played over just a month and there is much more scope for random variability in the performances of the teams. Thus, all in all, I think that their model is OK and will probably do well in terms of prediction. The nice feature is that they will update the predictions as more games are played (not sure they do it in a proper Bayesian fashion $-$ not saying they don't; just haven't read the details).


The predicted semi-finals are Spain-Germany and Italy-Netherlands. I think we'd be happy to reach the semis. We may have a shot, but I wouldn't be too surprised if we went back home much sooner (but that's me being a bit pessimistic, may be).

Monday, 4 June 2012

Norway

I vaguely knew about Norway, but having being there for a few days I could actually see and ask a few things about it. It's amazing and humbling how they managed to make the country work. The fact that they have invested 20 years on learning how to work with oil and gas (which is now paying off amazingly, with their sovereign wealth fund) is really a lesson that should be learnt by so many politicians (and voters!) around the world.


Surely it is much easier to run a country of 6million people than one that is 10 or 50 times as big, but still their welfare state is impressive.


On the other hand, the cost of living is ridiculously high: the equivalent of £5.03 for a coffee (check this out! $-$ as of just now, Google currency converter says £0.107 for 1 Norwegian Krone) really is steep. 


Anyway, very interesting place $-$ to visit.

Wednesday, 30 May 2012

Ready(?) for Trondheim

After the exam madness, I'm about to leave sunny England and off to relatively frosty Trondheim to attend and give a talk at the Second Workshop on Bayesian Inference for Latent Gaussian Models with Applications. Looks like a very interesting small conference, with lots of good talks.


Of course, I haven't finished preparing mine (but I'm unusually nearly there). I'll talk about my work on hierarchical models for data on In Vitro Fertilisation (IVF) and pre-implantation genetic screening (PGS). The main issue is that while we (or rather "they") are very good at fertilising eggs in a dish, we (again: they) are not so good at actually making the resulting embryo implant in the mother's womb, which means that the actual pregnancy rate is overall still about 30$-$40%.


The main problem is that once the eggs have been collected and fertilised, they are followed up for a few days after which the embryologist looks at them and decides which ones "look good"; these are then put back in the womb and hopefully develop in a pregnancy. But this of course does not account for potential problems in the genetic structure of the embryos, which are invisible to the naked eye. PGS is a complex genetic test that is able to detect potential chromosomal abnormalities in the embryos, thus making the choice of which are the "best" more informed. 


The theory that we are testing is that shorter telomere (some kind of protective cap sitting at the end of the chromosomes) are associated with chromosomal abnormalities. The nature of the data is obviously hierarchical, because normally we observe different cells, each of which belongs to one embryo, each of which is donated for research by a couple of parents. However, in this field there aren't many papers using appropriate methods to account for the implied correlation among the observations (in fact, I think there aren't any at all).


Our work seems to suggest that there is indeed some association between the length of telomere and chromosomal abnormalities $-$ especially for embryos that are not well developed. 




For "minor chaotic" embyros, which are as good as it gets in this setting, there doesn't seem to be any (or at least hardly any) effect of the length of telomere. In a way, even for the more problematic embryos (up to "uniformally abnormal", which are really screwed up) the effect is not very large and the probability of testing positive at PGS is not very much affected by the length of telomere. But for the intermediate embryos, the chance of testing positive to PGS is actually very much affected by telomere's length.


There are then potentially very important implications for the underlying science and for the management of IVF patients.


I'll post some more when I have finished the presentation. Also, we are working on a paper (just need to finalise a couple of details and re-submit it after having pleased the referees' thirst for modifications...)

Thursday, 24 May 2012

England is the new Spain

Pizza, gelato and motorbike with no jacket. 


Can the weather be any better? (as Chandler Bing would put it)

Tuesday, 22 May 2012

Eurovision again

Looks like I've not got many new topics these days, but here's an update on the Eurovision contest paper. 


Yesterday, we saw a shocking documentary on Panorama about the conditions in Azerbaijan (here, something related); specifically it showed how the "royal" family (which is not royal at all - technically they are a republic with free elections) are basically ruling the country as a proper family affair. Moreover, according to the report, they are using the contest (which they will host later this week as holders of the prestigious crown) as a wild propaganda tool.


Anyway, one of the most shocking bits was an interview to some guy (can't remember exactly who he was) who was arrested because he dared voting for Armenia, who are the historical enemy. He argued that he voted for them as a protest against the fact that state television blocked the broadcast when the Armenian act sung their song.


I went back to our preliminary analysis and produced this.



Azerbaijan is the country in white in the bottom right corner of the map. Darker colours indicate a higher propensity to vote for any of the other countries (ie from Azerbaijan to others), while lighter colours are an indication of lower propensity to do so.

You may recall that we are trying all sorts of models, accounting for "cultural" and "geographical" proximity - this one tries to include both. I think it's interesting that effectively all the countries in the former Soviet bloc (more or less all the rightmost part of the map) are coloured in solid dark grey, but Armenia (the light grey country bordering Azerbaijan).

The "effects" are pooled over the different clusters (regions/spatial proximity), which I think explains the fact that there is still some propensity for the voting pattern Azerbaijan $\Rightarrow$ Armenia. But the difference with all the other former Soviet bloc countries is quite large!

Bayes Pharma 2012 (again)

Julien has posted his impressions and comments on the conference on Christian Robert's blog


Agreed entirely, although, while I think INLA really is cool, I also believe that to have more than one software/algorithm available is absolutely a bonus - provided one know what one is doing (but again, see my comment somewhere else about the use of "one"). 


So, I don't think that we'll have problems of "duplications" with JAGS/BUGS or self-coded MCMC stuff!


PS: I love Julien's title - "Bayes on drugs"

Monday, 21 May 2012

Marking exam papers

I'm hopelessly in the dark land where you spend the whole of your time marking exam papers. And I burnt my tongue at lunch.


Doesn't get much worse than that...