Monday, September 17, 2012

TEAM A DAY 10: Avinash Pandey

CONJOINT ANALYSIS


Conjoint analysis is a statistical technique used in market research to determine how people value different features that make up an individual product or service.
The objective of conjoint analysis is to determine what combination of a limited number of attributes is most influential on respondent choice or decision making. A controlled set of potential products or services is shown to respondents and by analyzing how they make preferences between these products, the implicit valuation of the individual elements making up the product or service can be determined. These implicit valuations (utilities or part-worths) can be used to create market models that estimate market share, revenue and even profitability of new designs.
Conjoint analysis requires research participants to make a series of trade-offs. Analysis of these trade-offs will reveal the relative importance of component attributes. To improve the predictive ability of this analysis, research participants should be grouped into similar segments based on objectives, values and/or other factors.

First we listed attributes and asked respondents to give score and then some of all was given
the difference between the attributes is calculated and higher difference shows higher importance.
Coming to Generating orthogonal design
here we filled in all attributes when it came to a large number and cannot be done manually on excel, this takes in all attributes and finally gives us best 16 chosen lists which we can analyse on excel and calculate differences. It represents the best combinations.

Generate Orthogonal Design generates a data file containing an orthogonal main-effects design that permits the statistical testing of several factors without testing every combination of factor levels. This design can be displayed with the Display Design procedure, and the data file can be used by other procedures, such as Conjoint.
Example. A low-fare airline startup is interested in determining the relative importance to potential customers of the various factors that comprise its product offering. Price is clearly a primary factor, but how important are other factors, such as seat size, number of layovers, and whether or not a beverage/snack service is included? A survey asking respondents to rank product profiles representing all possible factor combinations is unreasonable given the large number of profiles. The Generate Orthogonal Design procedure creates a reduced set of product profiles that is small enough to include in a survey but large enough to assess the relative importance of each factor.

Day 11- Team F( Sidharth )



 We started off the 9 ‘o’ clock session with a very interesting topic ‘Conjoint Analysis’. It is one of the very important tools of Business Analytics and performs the analysis in a much better and concise manner. So what exactly is Conjoint Analysis? It is one of the most widely-used quantitative methods in Marketing Research. It is used to measure the perceived values of specific product features, to learn how demand for a particular product or service is related to price, and to forecast what the likely acceptance of a product would be if brought to market.
           

Respondents usually complete between 12 to 30 conjoint questions. The questions are designed carefully, using experimental design principles of independence and balance of the features. By independently varying the features that are shown to the respondents and observing the responses to the product profiles, the analyst can statistically deduce what product features are most desired and which attributes have the most impact on choice. In contrast to simpler survey research methods that directly ask respondents what they prefer or the important of each attribute, these preferences are derived from these relatively realistic tradeoff situations.
But the direct survey question "how much would you pay for xyz?" is unreliable and misleading. So instead, we ask the consumer's opinion on a series of products with differing features over a range of prices. Our techniques then use regression analysis to compute mathematical values that explain consumer behavior -  how much value is placed on price, or location, or features, etc. and then correlate this data to demographic, lifestyle, or other consumer profiles.





Decoding of the Conjoint Analysis Results:
Here’s an example of a survey regarding the consumer tastes in ice creams. The data is collated as per the consumer preferences.

Given the consumers' ratings of all 16 diverse combinations, the software package computes a mathematical regression to tell us how important each of the five factors is to the individual responding consumer, and to the group of responding consumers as a whole.
According to the results shown to the left (actual output from the online survey), we'd know that consumer X bases 47% of his decision on price, 23% on the flavor, 19% on the freshness, and is less concerned about the container or healthiness. We also learn get a relative ranking of the different flavors, as shown in the lower graph.


Maybe older customers who eat ice cream regularly are more concerned about healthiness.   Maybe younger consumers don't really care about the cone after all.  Perhaps those who work in a nearby office building and pass by for a snack really appreciate the homemade fresh ingredients.  All of these facts will be mathematically predicted using conjoint analysis. The end result is a quantitative, robust analysis of what consumers really want, with each attribute evaluated in the context of the others, incorporating the trade-offs that ultimately project the greatest influence on consumer behavior.

- By Sidharth 
Team F

Team I - day 10 -Rajdeep (Marketing)

Understanding consumer behavior using conjoint analysis using example

(note the responses generated for results are as per a study paper on the internet)


Conjoint Method

First, select what attributes of the product you would like to test, and what the possibilities are for each attribute. To demonstrate, let's use the example of an ice cream shop, which might want to know consumer attitudes about:

  • preferred flavor (vanilla, chocolate, strawberry, or black raspberry)
  • price ($1.50, $2.00, $2.50, $3.00)
  • container (cone, cup)
  • freshness (homemade & fresh, factory-produced)
  • healthiness (reduced fat, regular)

Is there a preferred flavor, or do customers like a variety? How much "extra" would someone be willing to pay more for a reduced-fat option? Do kids really prefer cones? How much do consumers value a neighborhood shop using fresh local ingredients?

The scientific way to answer these questions is to test each of the 5 attributes in the context of the others. To do that, we take each of these descriptors and create a series of "hypothetical" products, each with 5 attributes. The software creates templates for 16 (or 18 - depending upon the number of variables) of these, and we portray the description of the proposed product visually, on a "card", as shown to the right.

"Cards" can describe the product using words only - but can also use logos, pictures, or even smells or sounds. In any case, respondents will be asked to read each of the 16 "cards", and then assign a ranking of some kind (using numbers 1-x, or using adjectives like favorable, unfavorable, ideal, etc.)

Perhaps Card #1 is a factory-produced low-fat cheap vanilla cone. Maybe #2 is a homemade non-low-fat chocolate cup at a medium price point. The process goes on with 16 mathematically designed cards that offer all the relevant combinations of choices.
 

Conjoint Results
Given the consumers' ratings of all 16 diverse combinations, the software package computes a mathematical regression to tell us how important each of the five factors is to the individual responding consumer, and to the group of responding consumers as a whole.
According to the results shown to the left (actual output from the online survey), we'd know that consumer X bases 47% of his decision on price, 23% on the flavor, 19% on the freshness, and is less concerned about the container or healthiness. We also learn get a relative ranking of the different flavors, as shown in the lower graph.
In addition, each consumer will be asked a number of informational questions to create a demographic profile, so that we can compare the results and analyze them based upon income, age, location, and other variables that may affect consumer behavior towards a particular product.
Maybe older customers who eat ice cream regularly are more concerned about healthiness. Maybe younger consumers don't really care about the cone after all. Perhaps those who work in a nearby office building and pass by for a snack really appreciate the homemade fresh ingredients. All of these facts will be mathematically predicted using conjoint analysis.
The end result is a quantitative, robust analysis of what consumers really want, with each attribute evaluated in the context of the others, incorporating the trade-offs that ultimately project the greatest influence on consumer behavior.


 
 

DAY 10 Group I: Procedure to use of Conjoint in SPSS




Open SPSS Processor
Create plan file

































Create data in excel for analysis (convert excel file as per SPSS version. i.e SPSS version 15 can use only that excel file which are made in office 93-2003 format)


















Open excel file into SPSS processor
Save and close the excel file.
If the excel file is open during SPSS processing it will show error report.


















     
Select sheet from excel file


Save the current file, e.g  Data.sav
Open SPSS syntax file
Get syntax file from the location where  it is saved.


































   




You have to code syntax file as per given instructions.







While coding syntax file you need path of the every file created (i.e Data.sav, plan.sav,Data.exl etc)
Take curser to the file and open properties. Copy the path from the property box. Paste in the syntax box in front of respective instructions. Write name of files with their path. 





































Save syntax
Run Syntax file  (Run All)
Processor will generate the output





BY
Sushilkumar Balvir

Day 10-Team J-Manu Jain-SPSS Data Analysis Examples


SPSS Data Analysis Examples
Logit Regression

Version info: Code for this page was tested in SPSS 20.
Logistic regression, also called a logit model, is used to model dichotomous outcome variables. In the logit model the log odds of the outcome is modeled as a linear combination of the predictor variables.
Please note: The purpose of this page is to show how to use various data analysis commands. It does not cover all aspects of the research process which researchers are expected to do. In particular, it does not cover data cleaning and checking, verification of assumptions, model diagnostics and potential follow-up analyses.

Examples

Example 1:  Suppose that we are interested in the factors that influence whether a political candidate wins an election.  The outcome (response) variable is binary (0/1);  win or lose.  The predictor variables of interest are the amount of money spent on the campaign, the amount of time spent campaigning negatively and whether or not the candidate is an incumbent.
Example 2:  A researcher is interested in how variables, such as GRE (Graduate Record Exam scores), GPA (grade point average) and prestige of the undergraduate institution, effect admission into graduate school. The response variable, admit/don't admit, is a binary variable.

Description of the data

For our data analysis below, we are going to expand on Example 2 about getting into graduate school.  We have generated hypothetical data, which can be obtained from our website by clicking on binary.sav. You can store this anywhere you like, but the syntax below assumes it has been stored in the directory c:\data. This dataset has a binary response (outcome, dependent) variable called admit, which is equal to 1 if the individual was admitted to graduate school, and 0 otherwise. There are three predictor variables: gregpa, and rank. We will treat the variables gre and gpa as continuous. The variable rank takes on the values 1 through 4. Institutions with a rank of 1 have the highest prestige, while those with a rank of 4 have the lowest. We start out by opening the dataset and looking at some descriptive statistics.
get file = "c:\data\binary.sav".

descriptives /variables=gre gpa.

                                   

frequencies /variables = rank.


                      
                           

crosstabs /tables = admit by rank.


                                                          

  

Analysis methods you might consider

Below is a list of some analysis methods you may have encountered. Some of the methods listed are quite reasonable while others have either fallen out of favor or have limitations.
  • Logistic regression, the focus of this page.
  • Probit regression.  Probit analysis will produce results similar logistic regression. The choice of probit versus logit depends largely on individual preferences.
  • OLS regression.  When used with a binary response variable, this model is known as a linear probability model and can be used as a way to describe conditional probabilities. However, the errors (i.e., residuals) from the linear probability model violate the homoskedasticity and normality of errors assumptions of OLS regression, resulting in invalid standard errors and hypothesis tests. For a more thorough discussion of these and other problems with the linear probability model, see Long (1997, p. 38-40).
  • Two-group discriminant function analysis. A multivariate method for dichotomous outcome variables.
  • Hotelling's T2.  The 0/1 outcome is turned into the grouping variable, and the former predictors are turned into outcome variables. This will produce an overall test of significance but will not give individual coefficients for each variable, and it is unclear the extent to which each "predictor" is adjusted for the impact of the other "predictors."

Logistic regression

Below we use the logistic regression command to run a model predicting the outcome variable admit, using gregpa, and rank. The categorical option specifies that rank is a categorical rather than continuous variable. The output is shown in sections, each of which is discussed below.
logistic regression admit with gre gpa rank 
   /categorical = rank.
   
   

The first table above shows a breakdown of the number of cases used and not used in the analysis. The second table above gives the coding for the outcome variable, admit.
                                               

                                         
The table above shows how the values of the categorical variable rank were handled, there are terms (essentially dummy variables) in the model for rank=1, rank=2, and rank=3; rank=4 is the omitted category.


  • The first model in the output is a null model, that is, a model with no predictors.
  • The constant in the table labeled Variables in the Equation gives the unconditional log odds of admission (i.e., admit=1).
  • The table labeled Variables not in the Equation gives the results of a score test, also known as a Lagrange multiplier test. The column labeled Score gives the estimated change in model fit if the term is added to the model, the other two columns give the degrees of freedom, and p-value (labeled Sig.) for the estimated change. Based on the table above, all three of the predictors, gregpa, and rank, are expected to improve the fit of the model.


  • The first table above gives the overall test for the model that includes the predictors. The chi-square value of 41.46 with a p-value of less than 0.0005 tells us that our model as a whole fits significantly better than an empty model (i.e., a model with no predictors).
  • The -2*log likelihood (499.977) in the Model Summary table can be used in comparisons of nested models, but we won't show an example of that here. This table also gives two measures of psudeo R-square.


  • In the table labeled Variables in the Equation we see the coefficients, their standard errors, the Wald test statistic with associated  degrees of freedom and p-values, and the exponentiated coefficient (also known as an odds ratio).  Both gre and gpa are statistically significant. The overall (i.e., multiple degree of freedom) test for rank is given first, followed by the terms for rank=1, rank=2, and rank=3. The overall effect of rank is statistically significant, as are the terms for rank=1 and rank=2. The logistic regression coefficients give the change in the log odds of the outcome for a one unit increase in the predictor variable.
    • For every one unit change in gre, the log odds of admission (versus non-admission) increases by 0.002.
    • For a one unit increase in gpa, the log odds of being admitted to graduate school increases by 0.804.
    • The indicator variables for rank have a slightly different interpretation. For example, having attended an undergraduate institution with rank of 1, versus an institution with a rank of 4, increases the log odds of admission by 1.551.

Things to consider

  • Empty cells or small cells:  You should check for empty or small cells by doing a crosstab between categorical predictors and the outcome variable.  If a cell has very few cases (a small cell), the model may become unstable or it might not run at all.
  • Separation or quasi-separation (also called perfect prediction), a condition in which the outcome does not vary at some levels of the independent variables. See our page FAQ: What is complete or quasi-complete separation in logistic/probit regression and how do we deal with them? for information on models with perfect prediction.
  • Sample size:  Both logit and probit models require more cases than OLS regression because they use maximum likelihood estimation techniques.It is also important to keep in mind that when the outcome is rare, even if the overall dataset is large, it can be difficult to estimate a logit model.
  • Pseudo-R-squared:  Many different measures of psuedo-R-squared exist. They all attempt to provide information similar to that provided by R-squared in OLS regression; however, none of them can be interpreted exactly as R-squared in OLS regression is interpreted. For a discussion of various pseudo-R-squareds see Long and Freese (2006) or our FAQ page What are pseudo R-squareds?
  • Diagnostics:  The diagnostics for logistic regression are different from those for OLS regression. For a discussion of model diagnostics for logistic regression, see Hosmer and Lemeshow (2000, Chapter 5). Note that diagnostics done for logistic regression are similar to those done for probit regression.

Team C blog as on 17th September 2012 By Nitin Dhantole


BA as on 17th of September 2012:
Team C made by Nitin Dhantole

Revision in class was conducted and for my revision I have selected factor analysis for which I have taken sample data from internet on Film viewership.
I have used the data on factor analysis and got following results:
                                                Descriptive Statistics


Mean
Std. Deviation
Analysis N
Type of film viewed
2.06
.854
16
Age group of respondent
1.50
.516
16
Anxiety rating before watching film
4.56
1.861
16
Anxiety rating after watching film
5.63
2.604
16
Pulse rate before watching film
70.19
5.913
16
Pulse rate after watching film
73.06
6.913
16
Breathing rate before watching film
15.94
1.843
16
Breathing rate after watching film
17.13
2.363
16


                                      Communalities


Initial
Extraction
Type of film viewed
1.000
.940
Age group of respondent
1.000
.566
Anxiety rating before watching film
1.000
.949
Anxiety rating after watching film
1.000
.856
Pulse rate before watching film
1.000
.847
Pulse rate after watching film
1.000
.816
Breathing rate before watching film
1.000
.931
Breathing rate after watching film
1.000
.930
Extraction Method: Principal Component Analysis.

In this matrix all of the components are above 0.5 so I have selected all of them for my analysis.




This scree plot is perfect as per sir’s definition.
                                             Component Matrix(a)

 


Component
1
2
3
Type of film viewed
.595
.729
-.230
Age group of respondent
-.168
.127
.722
Anxiety rating before watching film
.828
-.457
.234
Anxiety rating after watching film
.844
.267
.270
Pulse rate before watching film
.658
-.499
-.406
Pulse rate after watching film
.798
.187
-.380
Breathing rate before watching film
.828
-.439
.229
Breathing rate after watching film
.874
.323
.249
Extraction Method: Principal Component Analysis.
a  3 components extracted.

                                                                




Rotated Component Matrix (a)

 


Component
1
2
3
Type of film viewed
-.101
.946
.188
Age group of respondent
.030
-.031
-.751
Anxiety rating before watching film
.951
.200
.067
Anxiety rating after watching film
.523
.753
-.126
Pulse rate before watching film
.652
.070
.645
Pulse rate after watching film
.334
.677
.497
Breathing rate before watching film
.938
.215
.068
Breathing rate after watching film
.503
.815
-.112
Extraction Method: Principal Component Analysis.
 Rotation Method: Varimax with Kaiser Normalization.
a  Rotation converged in 5 iterations.

From rotated component matrix I have go two components as liable which was labeled in excel as follows:
Component
Reaction of people before watching film
Reaction of people after watching film
3
Type of film viewed
-0.101
0.946
0.188
Age group of respondent
0.030
-0.031
-0.751
Anxiety rating before watching film
0.951
0.200
0.067
Anxiety rating after watching film
0.523
0.753
-0.126
Pulse rate before watching film
0.652
0.070
0.645
Pulse rate after watching film
0.334
0.677
0.497
Breathing rate before watching film
0.938
0.215
0.068
Breathing rate after watching film
0.503
0.815
-0.112
Extraction Method: Principal Component Analysis.
 Rotation Method: Varimax with Kaiser Normalization.
a
Rotation converged in 5 iterations.



Under this I have saved the components in the factor reduction save option and generated the graph of above two components and labeled as age to classify and generated following conclusion:


Conclusion:
As most of the people fewer than 18 are liable to react before watching a movie than the older ones.