Showing posts with label Philosophy. Show all posts
Showing posts with label Philosophy. Show all posts

Tuesday, 23 September 2014

Shathranj ke Khiladi (Experts of Chess)

TLDR;
Random musings about Chess and Life, The title of this post is from a story by Premchand, a Hindi writer who wrote pithy, wry and direct stories about aristocracy and the life of poor men under them. Experts of Chess briefly is a story of two chess players so engrossed in games of chess that they ignore the fall of their king and his kingdom and the decimation of soldiers, later they have an argument about the game and duel each other and are both mortally wounded which is the end of the story. This post has little to do with that story though and more about the game of chess and an analogy to life.

#include<chess.h>
#include<musings.h>
void main()
{
I’ve taken to playing chess about 4 months back. I was introduced by a friend to the game again (I used to play real lousy when I was a kid). She is good at it and has me beat all but 2 times out of the 12 games that we’ve played, there was a vodka fueled game where I was beaten really badly and I can’t still say it was because I was inebriated. Then I’ve played against the computer on Chess Titans on Windows 7. Started with Level 5, the first game lasted two hours with plenty of Undo Moves and a reversal from a checkmate (which according to the game statistics is counted as a loss even if you undo), played from there and won by checkmate. 

While I played I realized that early in the game there was no way to predict the future moves of my opponent to more than 3 moves as the number of possible combinations were insane. However, shit could get real very fast if you played fast and loose.

There is quite a bit of breathing room in the beginning and this is when you should prepare the layout of your pieces such that they back each other up. If you play a move just for the heck of it without an idea of what strategy that move benefits to, then you’ve wasted your move. A move that could tighten your offence or strengthen you defence is what you should always be on the lookout for.  Also no point taking out the big guns early when the board is full and you have no space to move. When the opponent is clever enough, even if you have them defended you will lose them to the pawns, which is never a good deal. Use the Knight early on, it can jump over pieces and on a choc-a-bloc board, it is a good way to pick out defensive pawn positions, avoid losing the Knight if you can help it. The computer always tries to get my Knights in exchange for its own and considers it a fair trade. I wonder why it would do that considering it could have used them just as well.

There are no spare moves towards the middle of the game. You have to play everything with a purpose and exactly right keeping all the pieces backing each other up while breaching into the enemy lines.  The middle of the game is when the pieces start going off the board and you have to start weighing your options. Would you mind losing a Knight for saving a rook or a bishop? A pawn threatens both your Knight and your bishop, which one do you choose? The Queen, yes you have to save her and you could end up sacrificing quite a bit just to save her. The computer always goes for your queen so never put it up-front with a lesser piece defending it. Send the Rook in for an attack if you have to and then defend the Rook with the Queen but never vice-versa because the computer will just take your Queen and not care what it loses in return. In other scenarios you will notice it putting its Queen head-to-head with yours and then losing its Queen to take yours. But this is a pre-meditated calculation that assumes that a Pawn from the computer is already half-way across the board and will reach the end of the board and promote itself to a Queen which will make your King a sitting duck.

The end of the game is a little anti-climactic. I’m told that towards the end, the game is predictable, so predictable that in fact the computer stores look-up tables for the moves that can result in the fastest check-mate. The irony is that, despite the board being largely empty, there is not much space to run around to and you will eventually be trapped into a corner and check-mated. Hope or horror also exists when a pawn goes over to the other side of the board and gets promoted, mostly to Queen. That is mostly the end, if you have just the King and other odd-end pawns, it could also be the end for the opponent if you have a bishop or another Queen on a largely empty board with the other King in the corner.

It seems to be quite a useful analogy for life. Play for a strong defense, build your fundamentals. When the attack comes you should have your defense ready otherwise it will be decimation. No point taking out the big guns first. Play the small battles with the pawns and weaken the defense of the enemy, then start moving in for the kill with the big guns

Recently I happened to play Level 6. It is very hard for me and this particular game lasted 2 hours or more. I got check-mated in at least 6 different ways and kept undoing the game to see what the point of no-return was. I played a variety of endings which were ending in check-mate. Finally I found an undo point where it seemed I was able to smuggle my Queen bang in the middle of enemy territory. It already had its Queen and kept troubling my King in an oscillating check series where I stayed put in a, for the lack of a better word, fortress of Pawns in a triangle. Then began my series of check moves, I had my Queen backed up with a Rook and that Rook backed up with another Rook just so that the computer Queen wouldn’t dare kill on the diagonal. Then a sudden check by the computer Queen which came and stood right behind my King and I took the Queen for nothing. Then there was nothing to do but to use my Queen as soon as possible for a check-mate and that was it. I had won Level 6 for the first time ever. I had drawn it previously but a win always has a nice warm fuzzy feeling to it. To beat a machine, well maybe not the best chess engine but a fairly decent one.

However if I were to stretch the analogy of chess for life far enough, is it possible that you could have made a series of moves from where there is a point of no return? It would not be really possible to know because I knew that until I did about 9 undoes (10 moves back more or less) I was not able to get to the point in the game where I was able to save my power pieces and also move in for the kill. I repeatedly lost my Rooks and the Queen in exchanges where the computer retained a Rook, lost the Queen and managed to promote a Pawn to Queen in the next 3 moves to give me a check-mate. Out of the 9 variants of the undo games I played (they all ended in check-mate for me differently) only one of the games was where I managed to defend as well as move in for the kill simultaneously.

Is it possible that I may have reached that point in life as well, where I may have just made the best of the positions I had set out initially but now it is just a few moves to check-mate? There is no Ctrl Z to the moves you make in life and there is of course no going back to play a different variant. Is the late 20’s, the middle game of academia-chess? Have the choices that determine the closing moves of the endgame been made? I guess I will know 5 years down the line and it will be too late to do anything by then. All I know is that going from one day to the next, I always make sure I always know more than what I started out with and knowledge has always been power. By a check-mate I mean the glass wall you hit up against and you just can't go higher than that, a win would be a place where you do the science you like and you live it too. Also when you realize that academic hierarchy like a lot of others in the world, is a little lop-sided, and you definitely don't want to be at the bottom of a pyramid of incompetents.

As to the question of if it ends in a check-mate for me, your guess is just as good as mine.

}

Saturday, 26 July 2014

Correlation: Philosophy and Practicality


This post deals with the philosophy of correlation and general statistical inference methods, my heavily biased opinions towards observational science and sickness associated with inferential methods. Then it is followed by some practical exercises on correlation, then some simulation follows and we try to see how robust correlation is to noise.

Correlation as I recently discovered is one of the corner-stone methods of statistical analysis as applied to biological sciences. Frequently I've come across papers that analyze high-throughput data to find a correlation between two apparently un-related sets of observations that were made and soon enough the paper is only filled with statements of correlation 0.6 P-value 10^-31 and so on. Towards the discussion, the authors try to explain some of the correlation in terms of some biological facts but I always had this feeling of dissatisfaction after reading such papers because I felt like I didn't really find out something amazing or new even after ploughing through it. To be honest I found such papers boring. It was more fascinating to read about  how they had tracked the motion of the ATP motor using a microscope than to read about hypothetical relations between sets of numbers obtained from experimental observation. It was only much later that I tried to find out what the whole deal with having a hypothesis was and why the question preceded the experimental techniques or the other way round, that I ended up finding out why.

Science, as it is done today follows two different philosophies. One is the old philosophy which had its roots in the way statistics was developed, called "hypothesis driven research". It involved posing a question initially which postulated an association between two previously-thought-to-be-un-related phenomena. This had to be done in quite a round-about way by framing it in the form of a "null hypothesis". So, if in your mind you had a feeling that the two phenomena were related, for the purposes of statistical analysis you have to think in the opposite way i.e. the phenomena are not related which becomes the null hypothesis (it isn't your hypothesis so you can't prove the null hypothesis; the theory that the phenomena aren't related). The alternate hypothesis is your hypothesis and essentially you have to come up with enough evidence to prove that the null hypothesis is incorrect. You cannot prove that your hypothesis (the alternative) is correct either. All you can do is prove that the null hypothesis is incorrect (with a certain probability that you might be wrong: the P-value) and that the two phenomena are related.

In my humble opinion, that is just a strange and weird way of doing it and unfortunately very confusing until you get used to it. In fact, the Bayesian school of thought teaches exactly the opposite. More on that much later. Now, the most common way to detect associations in statistics is the correlation metric.

The correlation metric measures the degree of linear association between two sets of numbers, in your case these two sets are experimental observations where one set is the numerical change in perturbation and the second set is the numerical output of the perturbed system or the observations. The value of correlation is scaled between -1 to 1 where numbers between 0~1 indicate positive correlation and between -1to 0 is a negative correlation. Positive correlation states that as the level of perturbation goes up the experimentally observed output of the system goes up. Negative correlation states the opposite.

Non-linear associations are not detected by the standard form of correlation that I'm discussing here called the Pearson correlation after Karl Pearson who created it. Non-linear associations exist in nature and their detection is done using other metrics like mutual information. An alternative way to detect non-linear trends using Pearson correlation is to convert all the numbers in your random variables X and Y to ranks and then do a Pearson correlation on these ranks. This alternative method of calculating correlation is called Spearman correlation. For linearly related numbers, Pearson and Spearman correlation will give the same value.



Practicals:

So correlation can be used to find if there is a linear trend between two sets of numbers. So if you were to construct a linearly related pair of number sets like

x<-seq(0,1,by=0.01)
y= 2*x + 40
plot(x,y,col="red")
cor(x,y)

Now even though the value of y is 2 times X and shifted by the number 40 it is still linearly related to the original set of numbers in X by the linear form (y=m.x+c). So the correlation is perfect and comes to a value of 1.

Now the numbers don't have to necessarily be in any sorted order to calculate correlation lest this example mislead you. You can re-order the numbers as long as you preserve the one-to-one relation between x and y  and you will get the same correlation value.

y_shuffle<-sample(y)
y_order<-match(y_shuffle, y)

So now that Y is shuffled but X is still the same you might get a terrible correlation (not 1) because the values have been all shuffled up. To check

cor(x,y_shuffle)

But if you shuffle X the same way you shuffled Y you can test the correlation between them and it should be the same as when the two sets were ordered.

x<-x[y_order]
cor(x,y)

That should have given you the same correlation as in the previous calculation.

Now correlation decreases when the points are moved away from the straight line that would pass through them. To see how that happens, you can try to scatter the points a bit by using some uniformly generated noise and add it to the data.

x<-seq(0,10,by=0.1)
y=3*x+20
cor(x,y)
# the correlation is 1

plot(x,y,type="b") # dead straight set of points
plot(x,y+runif(length(y),-0.2,0.2),col="purple",main="Scattering a line uniformly") # they get scattered a little randomly along the line



The correlation for this set of points is 0.9999173 which is really close to one, (your result will be different because we are using random numbers), by increasing the noise applied we can eventually destroy any semblance of correlation, but first we are going to explore this interesting concept I just thought of while writing now. What is the variation of the correlation when the same amount of noise is added to it. Since the noise we are adding is the same (-0.2, 0.2) we will get correlation values that are clustered closely, but what is the distribution of that clustering?

To check that, we will repeatedly calculate the correlation between x and y after scattering y with the same amount of uniformly distributed noise (-0.2, 0.2). After accumulating all these correlations into  a vector, we'll plot a histogram or density plot in this case.

cor_dist<-vector()
for(i in 1:5000){
 cor_dist<-c(cor_dist,cor(x,y+runif(length(y),-0.2,0.2)))
}

plot(density(cor_dist),type="h",col="red",main="Density plot of Correlation Scatter",xlab="Correlation Values")

Now I have a question, I added uniformly distributed random noise into the y axis, so shouldn't the values of correlation also fall uniformly between two intervals, why would I get a histogram where most of the correlation (1200) falls around one point. I think this has something to do with the magical properties of these statistics which try to converge to a certain distribution, I think they are called attractor distributions (ref: chaos theory) or stable distributions.

While one might be tempted to call this a bell curve or a normal distribution, it might be prudent to see the asymmetry on the right hand side of the distribution. Does that look like a slightly heavier tail? maybe.

You can test for normality using the Shapiro-Wilks test or the Kolmogorov-Smirnov test using.

shapiro.test(cor_dist)

One can also visualize normality using a qqplot which is a quantile-quantile plot that seeks to match a theoretical distribution against the distribution of the supplied data and find deviations from the straight line if any.

qqnorm(cor_dist,pch="*",col="purple")


Ok, we'll leave this for later, let's start to see how much is correlation destroyed with increasing amount of noise and scattering away from the straight line.

First we will do this the raw way, we'll add uniformly distributed noise to the intercept of our linear model and this noise will be added stochastically so the degradation of correlation may not be monotonic, in the next exercise we'll make it a little more deterministic by expressing the noise as a percentage of the original value. After some optimization I've realized that it is better to add the noise exponentially to be able to see how fast correlation degrades

x1=x

noise=2^(1:16) # we will use runif to add scatter symmetrically (negative  and positive) with number limits defined from this vector

noisyline<-sapply(noise,function(x){return(y+runif(length(x1),-(x/2),(x/2)))}) # So we pick exponential increasing ranges of value for noise and add it to the perfectly correlated variable and generate noisy Y values

noise_cor<-sapply(noise,function(x){return(cor(x1,y+runif(length(x1),-(x/2),(x/2))))}) # Now we calculate correlation for each x y set as y is destroyed by uniformly distributed noise that increases exponentially



#Plotting the files out to see how correlation changes with noise
# The scatterplot
for(i in 1:16){
png(paste("corrpoint",i,".png",sep=""))
plot(x1,noisyline[,i],col="blue",main="Noising up a Perfect Correlation")

dev.off()
 }
# The correlation value
for(i in 1:16){
png(paste("corrpoint",i,".png",sep=""),width=700) 
par(mfrow=c(1,2))
plot(x1,noisyline[,i],col="blue",main="Noising up a Perfect Correlation",xlab="X= 0 to 10,in steps of 0.01",ylab="Y= (3*X+20) + Noise")
plot(noise_cor[1:i],col="purple",main="Correlation decline on Noising",xlim=c(1,16),ylim=c(-0.5,1),type="b",xlab="Noising Iteration in Powers of 2",ylab="Correlation")
dev.off()}






So based on simulations, you can see a couple of things. For a small amount of noise the correlation metric does not fall drastically and then as the noise keeps increasing (exponentially) the correlation begins to fall in a steep slope to zero remaining robust up to 2^6 which is 64 which for y=3*x+20 where x=6 gives 38 which for the noise 64 added as x/2 gives a symmetrical -32 to +32 noise. So noising comparable to the values of the original numbers can result in correlation getting degraded which seems intuitive.

However what is interesting is when the points begin to scatter so much that it is pretty much a random scatter. At that point correlation begins to oscillate, 0.2 and lesser correlation can arise out of random scatter of perfectly correlated points to begin with. This calls for an introspection in terms of what value of correlation do we consider significant enough to point to a linear trend in the numbers that we are looking at. It is always good to eyeball the scatterplot rather than just blindly trusting the correlation value. For the purposes of publication however there are quantitative methods of establishing the significance testing for correlation. Significance is done using Permutation testing which we will look at in another post.


This post was long and it was just as hard for me to write it as it is for you to read. Take it slow and easy and try to understand what happens in the code and look at the graphs to see how you can also spot the correlation between a set of points by looking at how they are spread around the centre of the diagonal on a scatter-plot. When they are spread all over the place then the correlation is quite close to zero. However you will be surprised by how variations of the same number of data-points can give the same correlation or how certain regular shapes can have very little correlation. It sort of teaches you that a regularity or a symmetry in shape may not necessarily be grounds for a correlation but a symmetry around the diagonal of the scatterplot certainly is.