Tuesday, December 2, 2014

Random Forest using scikit learn in IPython

[RandomForestDemo]

Random Forest Classifier Demo

From scikit learn package import digits dataset
In [2]:
from sklearn.datasets import load_digits
In [3]:
digits_data = load_digits()
X = digits_data['data']
Y = digits_data['target'] 
In [4]:
X.shape
Out[4]:
(1797, 64)
In [5]:
Y
Y.shape
Out[5]:
(1797,)
In [6]:
import pylab as pl
pl.gray()
pl.matshow(digits_data.images[0])
pl.show()
<matplotlib.figure.Figure at 0xe483af0>
In [7]:
X[0]
Out[7]:
array([  0.,   0.,   5.,  13.,   9.,   1.,   0.,   0.,   0.,   0.,  13.,
        15.,  10.,  15.,   5.,   0.,   0.,   3.,  15.,   2.,   0.,  11.,
         8.,   0.,   0.,   4.,  12.,   0.,   0.,   8.,   8.,   0.,   0.,
         5.,   8.,   0.,   0.,   9.,   8.,   0.,   0.,   4.,  11.,   0.,
         1.,  12.,   7.,   0.,   0.,   2.,  14.,   5.,  10.,  12.,   0.,
         0.,   0.,   0.,   6.,  13.,  10.,   0.,   0.,   0.])
In [8]:
Y[0]
Out[8]:
0
In [9]:
pl.matshow(digits_data.images[1])
pl.show()
In [10]:
X[1]
Out[10]:
array([  0.,   0.,   0.,  12.,  13.,   5.,   0.,   0.,   0.,   0.,   0.,
        11.,  16.,   9.,   0.,   0.,   0.,   0.,   3.,  15.,  16.,   6.,
         0.,   0.,   0.,   7.,  15.,  16.,  16.,   2.,   0.,   0.,   0.,
         0.,   1.,  16.,  16.,   3.,   0.,   0.,   0.,   0.,   1.,  16.,
        16.,   6.,   0.,   0.,   0.,   0.,   1.,  16.,  16.,   6.,   0.,
         0.,   0.,   0.,   0.,  11.,  16.,  10.,   0.,   0.])
In [11]:
Y[1]
Out[11]:
1

From the corpus let us create Train and Test Dataset

In [12]:
from sklearn.cross_validation import train_test_split
In [13]:
x_train,x_test,y_train,y_test = train_test_split(X,Y,test_size=0.2,random_state=42)
In [14]:
x_train.shape
Out[14]:
(1437, 64)
In [15]:
x_test.shape
Out[15]:
(360, 64)
In [24]:
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score

clf = RandomForestClassifier(n_estimators=1,criterion="entropy")

clf.fit(x_train,y_train)
predictions = clf.predict(x_train)
In [25]:
predictions
Out[25]:
array([6, 0, 0, ..., 2, 7, 1])
In [26]:
print "Train Accuracy = %f "%(accuracy_score(y_train,predictions)*100)

predictions_test = clf.predict(x_test)

print "Test Accuracy = %f "%(accuracy_score(y_test,predictions_test)*100)
Train Accuracy = 92.414753 
Test Accuracy = 79.722222 

In [22]:
In []:

Tuesday, October 21, 2014

A cluster for Analytics

SetupAnaltyicsNode_blog

Analtyics Node

Recently I had setup a cluster for analytics purpose.

The machine has the following softwares installed.

  • Java 7
  • Base R
  • R-Studio
  • H2O Cluster

R Studio can be accessed through browser

http://<ip address>:8787/

In order to access H2O cluster, ssh tunnelling has be enabled from the client. Install H2O section has the details of tunnelling. Curently there are two

Install Java

sudo apt-get install python-software-properties
sudo add-apt-repository ppa:webupd8team/java
sudo apt-get update

Oracle JDK 7

sudo apt-get install oracle-java7-installer

Install R

sudo apt-key adv --keyserver keyserver.ubuntu.com --recv-keys E084DAB9 
echo "deb http://cran.cnr.berkeley.edu/bin/linux/ubuntu precise/" >> /etc/apt/sources.list
sudo apt-get update 
sudo apt-get install r-base

Type R to check installation

Install R Studio Server

64 bit version

$ sudo apt-get install gdebi-core
$ sudo apt-get install libapparmor1 # Required only for Ubuntu, not Debian
$ wget http://download2.rstudio.org/rstudio-server-0.98.994-amd64.deb
$ sudo gdebi rstudio-server-0.98.994-amd64.deb

Check by accessing http://:8787/ Use the linux username and password

Install H2o

Download H2o

http://s3.amazonaws.com/h2o-release/h2o/rel-kramer/1/index.html?smau_=iVV4RHKSHKjqQMBh

unzip h2o-2.4.6.1.zip`

Run h2o

java -Xmx2g -jar h20.jar
http://localhost:54322/

To access H2o from a remote machine. Setup a ssh tunnel in remote machine

ssh -L 55555:localhost:54321 tyconet@10.47.86.77

Access from remote machine using

http://localhost:55555

Install H2o cluster

Unzip h2o zip file in all the nodes of the cluster Create a nodes.txt file with the ip address of nodes, in this case it will only one entry

<ip address>:54321
<ip address>:54321

Start the H2o cluster in all the nodes

java -Xmx4g -jar h2o.jar -flatfile nodes.txt -port 54321

Tuesday, September 2, 2014

Platt scaling - calibrating classifier probabilities

Platts In many real world cases, it is very important to predict well calibrated probabilites.
Medical Domains, doctors will be more interested in the probability value associated with predictions, than a categorical output, which says if a patient has a certain condition or not.
In document clustering, a ranked affinity of documents to a cluster label will be more informative.
In this two part series, we explore how to go about calibrating output probabilities from classifiers. In this part we detail the problem and the use Reliability charts to visualize the output probabilities. Reliability charts indicates if there is a need to calibrate the probabilities.
Let us simulate some classification data for this exercise.
def load_data():
    X,Y = make_classification(n_samples = 100,n_features=100, random_state = 100)
    return X,Y
Use a maximum margin (SVM in this case) and Naive Bayes classifier
def max_margin_model(X,Y):
    mdl = SVC()
    mdl.fit(X,Y)
    return mdl

def nb_model(X,Y):
    mdl = BernoulliNB()
    mdl.fit(X,Y)
    return mdl
SVC does not returns the actual probability value, it returns the distance from the separating planes. We can scale these values between 0 and 1 to get the actual probability values.
We Fit the simulated data on these two classifiers and examine the output probability.
X,Y = load_data()
X_train,X_test,X_train,Y_test = train_test_split(X,Y,test_size=0.5)

mdl_nb = nb_model(X_train,Y_train)
predict_prob_nb = mdl_nb.predict_proba(X_test)

Y_t[Y==0]=-1

mdl_max_margin = max_margin_model(X_train,Y_train)
predict_prob_max_margin = mdl_max_margin.decision_function(X_test)
We need to look at the probability values of the true positives, so given the actual probability or a distance and true positive list, the following function returns the predicted probability for the true positives.
def getProbability(distances,Y,scale=False):
    Y_ones = numpy.where(Y == 1)[0]
    distances_ones = distances[Y_ones]
    if scale:
        min_max_scaler = preprocessing.MinMaxScaler()
        distances_ones = min_max_scaler.fit_transform(distances_ones)

    return distances_ones
For distance output from SVC, we set the scale variable to true, hence we get a number between 0 and 1. For naive bayes, we return the probability values for the true positives.
We use the above function to get the probabilty values attributed to the true positives, for both the models
predict_prob_nb = getProbability(predict_prob_nb[:,1],Y_test,False)
predict_prob_max_margin = getProbability(mdl_max_margin.decision_function(X_test),Y_test)
We look at the probability histogram,
bins = list(numpy.arange(0,1.1,0.1))
hist,bins = numpy.histogram(predict_prob_max_margin,bins)
>>> hist
array([ 9,  7, 12, 16, 19, 16, 24,  7, 21, 18])
Its interesting to see the spread of the probability values for SVC.
As its seen in the output values, SVM is trying to push the predicted output value away from 0 and 1.
With naive bayes
>>> hist,bins = numpy.histogram(predict_prob_nb,bins)
>>> hist
array([ 25,   1,   2,  10,   9,  25,  17,  15,  12, 131])
The probability values are pushed towards 0 or 1, its evident from the below plot.
As expected, naive bayes is pushing the probability values towards 0 and 1.

Reference

http://www.cs.cornell.edu/~alexn/papers/calibration.icml05.crc.rev3.pdf http://www.stat.cmu.edu/~fienberg/Statistics36-756/DegrootFienberg-Statistician-1983.pdf

Tuesday, August 19, 2014

HBase write throughput

Hbase_write_throughput

HBase write throughput as a function of number of column qualifiers

In Hbase, every cell value is stored along with all its cardinalities as follows,

rowkey:columnfamily:columnqualifier:timestamp:value

Hypothetically, let us assume the following

Data payload size               = 10 kb
rowkey size                     = 64 kb
columnfamily:columnname size    = 60 kb

In order to write a row with say 2 columns, the total amount of bytes transferred and written will be

2 * ( 5kb + 64 kb + 60 kb) = 258 kb (Total 10kb of payload split between two columns)

In order to write a row with say 1 column, the total will be

1 * (10 + 64 + 60) = 134 kb.

Larger the size, more data transfer across network, memstore will get full more often and hence will need more flush. This will negatively impact write throughput.

Verfiying this behaviour using HBase Load Testing tool,

Summary

- Rows          :   10k             10k             10K
- Columns       :   2               5               10
- PayLoad       :   512 kb          200 kb          100 kb
- Total PayLoad :   ~1000 kb        1000 kb         1000 kb     
- Throughput    :   405 Keys/s      252 keys/s      175 keys/s

Details

We start with 10000 rows, 2 columns with a payload of 512kb for every cell, indicated by - write 2:512:20

$ hbase org.apache.hadoop.hbase.util.LoadTestTool -write 2:512:20 -num_keys 10000
14/08/19 11:47:26 INFO Configuration.deprecation: hadoop.native.lib is deprecated. Instead, use io.native.lib.available
Key range: [0..9999]
Multi-puts: false
Columns per key: 1..4
Data size per column: 256..768

Below is a log captured at 5 seconds interval, at the end of 20 seconds, we see that write throughput is 405 keys/s

Starting to write data...
14/08/19 11:47:39 INFO util.MultiThreadedAction: [W:20] Keys=1663, cols=5.6 K, time=00:00:05 Overall: [keys/s= 332, latency=58 ms] Current: [keys/s=332, latency=58 ms], wroteUpTo=-1
14/08/19 11:47:44 INFO util.MultiThreadedAction: [W:20] Keys=3641, cols=12.3 K, time=00:00:10 Overall: [keys/s= 361, latency=54 ms] Current: [keys/s=395, latency=51 ms], wroteUpTo=-1
14/08/19 11:47:49 INFO util.MultiThreadedAction: [W:20] Keys=5769, cols=19.5 K, time=00:00:15 Overall: [keys/s= 382, latency=51 ms] Current: [keys/s=425, latency=46 ms], wroteUpTo=-1
14/08/19 11:47:54 INFO util.MultiThreadedAction: [W:20] Keys=8128, cols=27.6 K, time=00:00:20 Overall: [keys/s= 405, latency=49 ms] Current: [keys/s=471, latency=42 ms], wroteUpTo=-1
Failed to write keys: 0

We do it again with 5 columns, 200kb payload and 10k rows

$ hbase org.apache.hadoop.hbase.util.LoadTestTool -write 5:200:20 -num_keys 10000
14/08/19 14:38:20 INFO Configuration.deprecation: hadoop.native.lib is deprecated. Instead, use io.native.lib.available
Key range: [0..9999]
Multi-puts: false
Columns per key: 1..10
Data size per column: 100..300
.
.
Starting to write data...
14/08/19 14:38:31 INFO util.MultiThreadedAction: [W:20] Keys=901, cols=5.6 K, time=00:00:05 Overall: [keys/s= 180, latency=106 ms] Current: [keys/s=180, latency=106 ms], wroteUpTo=-1
14/08/19 14:38:36 INFO util.MultiThreadedAction: [W:20] Keys=1979, cols=12.4 K, time=00:00:10 Overall: [keys/s= 197, latency=99 ms] Current: [keys/s=215, latency=92 ms], wroteUpTo=-1
14/08/19 14:38:41 INFO util.MultiThreadedAction: [W:20] Keys=3070, cols=19.3 K, time=00:00:15 Overall: [keys/s= 204, latency=96 ms] Current: [keys/s=218, latency=91 ms], wroteUpTo=-1
14/08/19 14:38:46 INFO util.MultiThreadedAction: [W:20] Keys=4367, cols=27.7 K, time=00:00:20 Overall: [keys/s= 218, latency=90 ms] Current: [keys/s=259, latency=77 ms], wroteUpTo=-1
14/08/19 14:38:51 INFO util.MultiThreadedAction: [W:20] Keys=5857, cols=36.9 K, time=00:00:25 Overall: [keys/s= 234, latency=84 ms] Current: [keys/s=298, latency=66 ms], wroteUpTo=-1
14/08/19 14:38:56 INFO util.MultiThreadedAction: [W:20] Keys=7373, cols=46.4 K, time=00:00:30 Overall: [keys/s= 245, latency=80 ms] Current: [keys/s=303, latency=65 ms], wroteUpTo=-1
14/08/19 14:39:01 INFO util.MultiThreadedAction: [W:20] Keys=8843, cols=55.7 K, time=00:00:35 Overall: [keys/s= 252, latency=78 ms] Current: [keys/s=294, latency=67 ms], wroteUpTo=-1
Failed to write keys: 0

As seen above, the write throughput has reduced to 252 keys/s.

Further increasing the number of columns to 10, with 100K payload, the write throughput is reduced to 175 keys/s

$ hbase org.apache.hadoop.hbase.util.LoadTestTool -write 10:100:20 -num_keys 10000
14/08/19 14:34:54 INFO Configuration.deprecation: hadoop.native.lib is deprecated. Instead, use io.native.lib.available
Key range: [0..9999]
Multi-puts: false
Columns per key: 1..20
Data size per column: 50..150

Starting to write data...
14/08/19 14:35:07 INFO util.MultiThreadedAction: [W:20] Keys=582, cols=6.4 K, time=00:00:05 Overall: [keys/s= 116, latency=168 ms] Current: [keys/s=116, latency=168 ms], wroteUpTo=-1
14/08/19 14:35:12 INFO util.MultiThreadedAction: [W:20] Keys=1157, cols=13.0 K, time=00:00:10 Overall: [keys/s= 115, latency=171 ms] Current: [keys/s=115, latency=173 ms], wroteUpTo=-1
14/08/19 14:35:17 INFO util.MultiThreadedAction: [W:20] Keys=1884, cols=21.0 K, time=00:00:15 Overall: [keys/s= 125, latency=158 ms] Current: [keys/s=145, latency=137 ms], wroteUpTo=-1
14/08/19 14:35:22 INFO util.MultiThreadedAction: [W:20] Keys=2687, cols=30.0 K, time=00:00:20 Overall: [keys/s= 134, latency=147 ms] Current: [keys/s=160, latency=123 ms], wroteUpTo=-1
14/08/19 14:35:27 INFO util.MultiThreadedAction: [W:20] Keys=3558, cols=39.8 K, time=00:00:25 Overall: [keys/s= 142, latency=139 ms] Current: [keys/s=174, latency=115 ms], wroteUpTo=-1
14/08/19 14:35:32 INFO util.MultiThreadedAction: [W:20] Keys=4513, cols=50.5 K, time=00:00:30 Overall: [keys/s= 150, latency=132 ms] Current: [keys/s=191, latency=104 ms], wroteUpTo=-1
14/08/19 14:35:37 INFO util.MultiThreadedAction: [W:20] Keys=5410, cols=60.5 K, time=00:00:35 Overall: [keys/s= 154, latency=128 ms] Current: [keys/s=179, latency=111 ms], wroteUpTo=-1
14/08/19 14:35:42 INFO util.MultiThreadedAction: [W:20] Keys=6322, cols=70.8 K, time=00:00:40 Overall: [keys/s= 157, latency=126 ms] Current: [keys/s=182, latency=109 ms], wroteUpTo=-1
14/08/19 14:35:47 INFO util.MultiThreadedAction: [W:20] Keys=7280, cols=81.8 K, time=00:00:45 Overall: [keys/s= 161, latency=123 ms] Current: [keys/s=191, latency=104 ms], wroteUpTo=-1
14/08/19 14:35:52 INFO util.MultiThreadedAction: [W:20] Keys=8496, cols=95.6 K, time=00:00:50 Overall: [keys/s= 169, latency=117 ms] Current: [keys/s=243, latency=82 ms], wroteUpTo=-1
14/08/19 14:35:57 INFO util.MultiThreadedAction: [W:20] Keys=9632, cols=108.1 K, time=00:00:55 Overall: [keys/s= 175, latency=113 ms] Current: [keys/s=227, latency=87 ms], wroteUpTo=-1
Failed to write keys: 0

Monday, August 11, 2014

Unsupervised Feature Learning - Sparse Filtering

SparseFiltering

Sparse Filtering

An unsupervised feature learning method. Click here for the paper.
The code is available in mygithub
A small python implementation of the Sparse Filtering algorithm. This implementation has dependency on
As discussed in the paper we use a soft absolute activation function.
def soft_absolute(v):
    return np.sqrt(v**2 + epsilon)
The input X dimension is (nsamples,ndimensions). For testing purpose, we use scikit learn's make_classification function. We create 500 samples and 100 features.
def load_data():
    X,Y = make_classification(n_samples = 500,n_features=100)
    return X,Y
We use a simple Support Vector Machine Linear classifier to do final classiciation.
def simple_model(X,Y):
    clf_org_x = SVC()
    clf_org_x.fit(X,Y)
    predict = clf_org_x.predict(X)
    acc=  accuracy_score(Y,predict)
    return acc
We train a two layer network.
X,Y = load_data()
acc = simple_model(X,Y)

X_trans = sfiltering(X,25)

acc1= simple_model(X_trans,Y)

X_trans1 = sfiltering(X_trans,10)

acc2= simple_model(X_trans1,Y)

print "Without sparsefiltering, accuracy = %f "%(acc)
print "One Layer Accuracy, = %f, Increase = %f"%(acc1,acc1-acc)
print "Two Layer Accuracy,  = %f, Increase = %f"%(acc2,acc2-acc1)
At the first layer, we create 25 features. At the second layer we reduce them to 10. Finally a (500,10) X matrix is used by the SVC classifier.
Without sparsefiltering, accuracy = 0.986000 
One Layer Accuracy, = 1.000000, Increase = 0.014000
Two Layer Accuracy,  = 1.000000, Increase = 0.000000
With a single layer sparse filtering the accuracy reaches 100%. The second layer is redundant here.
Other implementations available in the web are,

Saturday, August 2, 2014

Unsupervised key phrase extraction from document - RAKE

RAKE

RAKE is an extremely effiecient keyword extraction algorithm and operates on individual documents. Its language and domain independent.

Literature is abundant with methods which uses Noun phrase chunks,POS tags, ngram statistics and similar others.

Given a document, stop word list and a list of phrase delimiters, RAKE extracts candidate phrases.

phrases = []

for line in sentences:
    words = nltk.word_tokenize(line)
    phrase = ''
    for word in words:
        if word not in stopwords.words('english') and word not in [',','.','?',':',';']:
            phrase+=word + ' '
        else:
            if phrase != '':
            phrases.append(phrase.strip())
            phrase = ''

We takes the stop word list from nltk and use a list of special characters as word delimiters.Every sentence is split into chunks at stop words occurence and phrase delimiters occurence.

The next step is to find out the frequency of the individual words in these phrases. Frequency is the count of occurence of the word.

word_freq = defaultdict(int)
word_degree = defaultdict(int)
word_score = defaultdict(float)

for phrase in phrases:
    words = phrase.split(' ')
    phrase_length = len(words)
    for word in words:
        word_freq[word]+=1
        word_degree[word]+=phrase_length

Degree of a word is sum of length of all the phrases where the word occurs. Finally scoring for each word is done by

score(word) = degree(word) / freq(word).

for word,freq in word_freq.items():
    degree = word_degree[word]
    score = ( 1.0 * degree ) / (1.0 * freq )
    word_score[word] = score

Phrase scores are calcuated by sum of individual word scores in that phrase. Phrase with very large scores are considered to be keyphrases for the document.

RAKE algorithm is explained in the book Text Mining:Application and Theory

There are several python implementations available.

The current code is available in github