MKF2121 Chap.5 Association, Regression, and Segmentation
Association, Regression, and Segmentation
Association analysis connects variable type to cross-tabulation, chi-square, correlation and regression. This chapter emphasises conditional percentages, effect magnitude, coefficient units, residual diagnosis, prediction validation and the boundary between a useful predictor and a proven causal lever. Segments are judged by distinctiveness, reachability, stability and actionability.
What this chapter covers
- 01
Variable types and association routes
- 02
Cross-tab counts and conditional percentages
- 03
Chi-square expected counts and effect size
- 04
Pearson correlation direction and strength
- 05
Regression intercepts, slopes and residuals
- 06
F tests, coefficient tests and R-squared
- 07
Multiple regression and conditional interpretation
- 08
Training, validation and test roles
- 09
Segmentation profiles and solution stability
- 10
From statistical pattern to testable marketing action
Interpret regression anatomy correctly
- 2Use the F test for the predictor model versus an intercept-only model.
- 2Interpret R-squared as fitted sample variation and b as outcome change per predictor unit.
- 2Check residuals, predictor range, new-data performance and the non-causal design boundary.
Key terms
- Cross-tabulation
- A table of joint counts or conditional percentages for categorical variables.
- Chi-square test
- A test comparing observed cell counts with counts expected under independence.
- Pearson correlation
- A coefficient describing the direction and strength of linear association between metric variables.
- Regression intercept
- The fitted outcome value when every predictor in the equation equals zero.
- Regression slope
- The expected outcome change associated with one predictor unit under the model.
- Market segment
- A customer group with a meaningful profile that can support differentiated marketing action.
Association, Regression, and Segmentation FAQ
Which percentage direction should a cross-tab use?
Condition on the groups named in the research comparison. If the question asks channel preference within segments, each segment is the denominator. State within before the percentage so the interpretation cannot silently reverse.
What does R-squared tell a marketing researcher?
It describes the share of observed outcome variation fitted by the model in the analysis sample. It does not show the percentage of cases predicted correctly, causal strength or performance on future data.
When is a statistical segment useful?
A segment must have a distinct and relevant profile, sufficient opportunity, an ethical route for identification and contact, stability across plausible samples or model choices, and a strategy that genuinely differs from other groups.
Assessment move
Practise cross-tabs by writing the denominator before calculating percentages, then interpret chi-square together with the cell pattern and Phi or Cramer's V. For correlation, always inspect direction, form, range and influential cases. For regression, label intercept, slope, residual, F test and R-squared separately, and state units and held-constant variables.
Finish every association with the causal boundary and a next design. Evaluate segments by reachability and actionability rather than by mathematical separation alone. Build cross-tab drills in which the same counts answer two conditional questions. Calculate percentages within rows and within columns, label each denominator and explain why the interpretations differ.
Derive expected counts from the margins, identify cells contributing to chi-square and pair the significance decision with Phi or Cramer's V. For correlation, sketch positive, negative, curved, clustered and influential-point patterns that could produce misleading summaries. State scale direction before interpreting the sign.
For regression, annotate an equation with outcome units, predictor units, intercept reference and slope meaning. Use a plotted observed range to mark unsupported extrapolation. Separate the omnibus F question, the conditional coefficient question and the R-squared fit description. Add residual diagnostics and one omitted-variable story before writing a causal sentence.
If the aim is prediction, define the prediction time and split training, validation and test data without leaking future records or repeated customers. Compare performance with a simple benchmark and inspect errors across relevant groups. For segmentation, evaluate alternative solutions using distinctiveness, size, reachability, stability and actionability.
Replace catchy labels with evidence-based profiles and describe how assignment would work in practice. End each session with a decision chain: observed pattern, magnitude, represented population, strongest permitted inference, proposed action and a design that can test whether the action changes the outcome. This closing move turns association into responsible marketing learning rather than causal storytelling.
Run a transfer challenge on the final model or segment. Change the population, period, predictor range or implementation and ask which conclusion survives. Separate stable descriptive evidence from assumptions that need validation. Then specify a feasible intervention, an outcome window and a comparison that can test it.
This keeps association useful for discovery without presenting the fitted relationship as a permanent causal rule.
Working through Association, Regression, and Segmentation in MKF2121? Sia is AskSia’s AI Marketing tutor — ask any MKF2121 Association, Regression, and Segmentation question and get a clear, step-by-step explanation grounded in how MKF2121 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.