Showing posts with label data mining business intelligence class Fayetteville Arkansas. Show all posts
Showing posts with label data mining business intelligence class Fayetteville Arkansas. Show all posts

Tuesday, February 14, 2012

Prediction with logistic regression

Yesterday evening we had the 3rd meeting of our. The tradition here at the University of Arkansas is to start with linear regression and logistic regression. Then go on with decision tree and other algorithms. I skipped linear regression, explained logistic regression theory and moved on to dissecting and interpreting SAS EM results. The data set is telecom Churn data set which is available on the book website, has over 3000 records and is very clean.


The ppt slides can be found here. I started with a simple example of linear regression: Gorgiean's enjoyment of snow over time which I found from here. Interestingly we had snow on Monday, there was some snow on the ground when we woke up. Thus this example was very relevant !!! My students liked it.
I then talked about data partitioning. I explained the reason for moving from:


target -> probability of target -> odds of target -> log of odds -> conducting linear regression of log of odds on the input variables.


It is not easy for an undergraduate student to digest these but I wanted to expose them to the ideas. I recommend this to all teachers to spend some time on explaining the assumptions and theory behind logistic regression.


On SAS EM output, I focused on coefficient estimates, significance levels for each input and for the whole regression equation, misclassification rate, false positives, false negatives, lift, and lift chart. The class activity and the follow up concept checks are available here.


I plan to talk about stepwise, forward, and backward variable selection methods in next class- and move on to KNN.

Monday, February 6, 2012

Touching the data!

Today was the second meeting of our BI class. Students had done a simple exercise in SAS EM with Churn data set which has 3351 data records. It's the same telecom Churn data I had used for my professional Master's class but it's clean.


The topic was data preprocessing and some basic EDA. We talked about correlated variables, Chi-square, Cramer's V, variable worth, missing values, normalization, outliers, mean, mode, median, skewness, kurtosis. The PPT slides are here. Student reminded me that I needed to explain what interval, nominal, ordinal, and binary variable are, it was a timely reminder. Phone numbers, area code, and zipcode are always good examples. 


Class activity for today's class is here, so is the concept check which is on the second page on the same link.

Monday, January 30, 2012

First meeting of my Undergraduate Business Intelligence Course

Yesterday, Monday, Jan. 30, 2012 was the first meeting of my ISYS 4293, Business Intelligence class at the University of Arkansas, Fayetteville. 


Our BI course is focused on Data Mining and we use Daniel Larose's book, Discovering Knowledge in Data (2005). Here are my course syllabus and schedule.


Here are my PowerPoint Slides for the first class.  I had asked students to read a piece of news related to DM on KDnuggets.com or other websites and share it with the class. I had also asked the class to review previous year's KDD Cup and choose a data set to work on during the class. The class was pretty excited. They had read the news and had researched the KDD Cup.They, however, know little about DM and thus some of the evaluations measures such as AUC and RMSE became confusing. We covered the basics, CRISP-DM, and DM tasks, discussed the 5 case studies in the book and as I walked around the class we discussed some of the KDD Cup data sets. 


I am excited, tonight working on Assignment 2 for the class, would want the first assignment covers all the details the class needs to learn from Chapter 2.