STAT 432
  • Welcome
  • Lectures
    • Overview
    • Week 1: Setup and AI Tools
    • Week 2: Training and Test Error
    • Week 3: Ridge Regression and Optimization
    • Week 4: Lasso and Variable Selection
    • Week 5: K-Nearest Neighbors
    • Week 6: Classification Error and Evaluation
  • Discussion
  • Quizzes
  • Final Project
  • Syllabus
  • Canvas
Skip to main content

Week 2 Discussion Questions

Selected STAT 432 Week 2 questions on model selection, prediction error, cross-validation, and AI workflows.
Back to weekly question pools and results

Week 2 results

Discussion honors

Team honors are based on the class ratings recorded during the discussion.

  1. 1

    Gold

    Team 2

    Team members

    • Haoran
    • Shih-Peng
    • Timothy
    • Ting-Yun
    • Zexuan
  2. 2

    Silver

    Team 1

    Team members

    • Brisa
    • Jessica
    • Lyndsey
    • Maeve
    • Ricky
  3. 3

    Bronze

    Team 3

    Team members

    • Heta
    • Jingyi
    • Mann
    • Neel
    • Zhijing
C

Top challenger

  • Lu Gong

These questions were selected and adapted from the Week 2 student submissions. Work through the reasoning with your group and be prepared to explain your conclusions.

Eleven questions are listed below. Contributor NetIDs appear at the end of each question. Questions discussed in class are marked.

Question 01Model selection criteria

When AIC, BIC, and Cp Disagree (Discussed)

Suppose we compare candidate linear regression models fitted to the same dataset. Under what conditions might AIC, BIC, and Mallows’ CpC_p select different models? Describe a situation in which the improvement in fit is enough to justify an additional predictor under one criterion but not another. Explain how the fit term, the complexity penalty, and the sample size enter that decision.

This question is contributed by brianc19, owenp3, and willm4.

Question 02Model selection criteria

Why BIC Often Selects a Smaller Model (Discussed)

Why does BIC usually favor smaller linear regression models than AIC or Mallows’ CpC_p? Under what modeling goals and assumptions would you prefer BIC? Explain why choosing the smaller model is not automatically the better decision, even when the candidate models have similar validation MSE.

This question is contributed by gjia3, hx27, and schuang7.

Question 03Prediction and explanation

Prediction or Explanation: Two Real Decisions (Discussed)

Give two specific real-life examples of using a regression model: one where accurate prediction is the main goal, and another where understanding the relationship between particular predictors and the response is the main goal. For each example, state the decision the analysis should support and explain how that goal affects which model you would prefer. Why might the model with the lowest prediction error be less useful for the second task?

This question is contributed by haorans9 and ousher2.

Question 04Prediction error

When Expected Test Error Is Not U-Shaped (Discussed)

Consider a sequence of nested least-squares models that add predictors one at a time, with the training sample size and design held fixed as in the Week 2 lecture. Must expected test MSE have a U-shaped curve as the number of predictors increases? Give two specific examples in which the theoretical curve is not U-shaped over the range of candidate models being compared. Explain each example using approximation bias and estimation variance, rather than fluctuations from one simulation run.

This question is adapted from a related question contributed by stewary2.

Question 05Prediction targets

Predicting for One Target or an Entire Population

Give a real-life example in which the best model for predicting the response of someone or something with a particular set of characteristics could differ from the best model for minimizing average prediction error across the whole population. Describe the specific target, the population, and a predictor whose usefulness differs between the two tasks. Explain in words why including that predictor could help one task more than the other. Use no code or mathematical formulas.

This question is contributed by cudzich3, gaoxinc2, and mannat2.

Question 06Model search

Best Subset Selection Versus Stepwise Selection (Discussed)

Compare best subset selection with stepwise selection when both use the same model-selection criterion. What are the advantages and disadvantages of each in terms of computation, the models searched, and dependence on earlier selection decisions? Describe a situation in which stepwise selection could miss a model found by best subset selection. Does finding the smallest criterion value among all subsets guarantee the lowest error on future data?

This question is contributed by chahak2 and ruitong7.

Question 07Cross-validation

Why Variable Selection Must Be Inside Cross-Validation

An analyst uses the entire dataset to select predictors, then refits that selected model separately in each training fold of 10-fold cross-validation. The analyst argues that the resulting MSE is an honest estimate of future prediction error because every observation is held out once. What information has already leaked into the validation folds? Describe how to evaluate the full selection-and-fitting procedure correctly, including where predictor selection happens and when a separate final test set may be used.

This question is contributed by nnigam2, shushim2, and wenhao7.

Question 08Prediction error

When One Test Set Favors the Worse Model on Average

Suppose adding a predictor increases expected test MSE, averaged over repeated training and independent test responses as in the Week 2 lecture. Could the observed test MSE nevertheless decrease on one particular training/test response pair? Explain what varies from one repetition to another and why a single observed improvement does not contradict the expected-error comparison.

This question is contributed by tingyun3.

Question 09Prediction uncertainty

Why More Data Cannot Eliminate Prediction Uncertainty

At a fixed set of predictor values, suppose the response has nonzero irreducible noise variance. A student claims that with enough training data, a 95% prediction interval for one new response can become arbitrarily narrow, just like a 95% confidence interval for the mean response. Where does this reasoning break down? Explain which source of uncertainty can shrink as the training sample grows and which remains for the new response.

This question is contributed by yanxif2.

Question 10AI workflow: data paths

An AI Agent Finds the Wrong Copy of the Data

While helping with Homework 2, an AI agent writes code that searches several folders for diabetes.csv and loads the first copy it finds. Suppose the search succeeds, but it finds an older copy outside the intended course data folder. What should you ask the agent to check before trusting the analysis, and how should it revise the file-loading code so another student can tell which dataset will be used when running the repository on a different computer?

This question is contributed by jyan36.

Question 11AI workflow: error diagnosis

Should an AI Agent Rewrite Code After a PermissionError? (Discussed)

An AI agent runs a Python Quarto document, but rendering stops with a PermissionError when Quarto tries to write a log file. The agent proposes changing the regression code. What evidence in the error message and traceback should it examine before editing anything? Describe the next diagnostic step you would ask it to take and how the result would help distinguish a problem with the statistical code from a problem with the output location or required access.

This question is contributed by ziqiz13.

STAT 432 | Basics of Statistical Learning

 
  • Instructor