RMIT University · FACULTY OF DATA SCIENCE

COSC2670 Chap.5 Classification Neighbours and Decision Trees

- one subject, every graph, every model, every mark
5 Chapters2-page Bible
Our own words - no uploaded lecturer files
Updated for this semester
Chapter 5 of 7 · COSC2670

Classification Neighbours and Decision Trees

Define classification

The course material gives this chapter a concrete anchor: The classification block develops k-nearest neighbours, distance choices, confusion matrices, decision trees, Gini splitting and cross-validation.

That classification anchor controls how nearest-neighbour rule is explained and how decision tree is tested in changed practice.

Classification Neighbours and Decision Trees is a quantitative decision problem built from classification, nearest-neighbour rule and decision tree.

The aim is to compare a distance-based classifier with a tree while controlling feature scale, depth and validation; a numerical result earns meaning only when the variables, units, assumptions and comparison are all explicit.

Begin with classification: state what quantity it represents, the scale on which it is measured and the condition under which it changes.

Then map every symbol in the Classification Neighbours and Decision Trees formula checkpoint to classification before calculation begins.

Next connect nearest-neighbour rule to the calculation. Show the nearest-neighbour rule transformation line by line, preserve units and signs, and make any denominator or baseline visible.

A nearest-neighbour rule calculator output is not a method; the reader must be able to reconstruct why that operation answers the question.

Formula checkpoint: classification

Euclidean distance
d(x,z)=j=1p(xjzj)2d(x,z)=\sqrt{\sum_{j=1}^{p}(x_j-z_j)^2}

Distance aggregates feature differences, so scale and irrelevant dimensions can dominate the neighbour relation unless preprocessing is justified.

Trace nearest-neighbour rule

Use decision tree to interpret or stress-test the result.

Ask whether the decision tree magnitude is plausible, whether a boundary case behaves as expected and which conclusion would reverse if an assumption changed. This is where computation becomes analysis rather than arithmetic.

When the task is to compare a distance-based classifier with a tree while controlling feature scale, depth and validation, separate inputs supplied by the problem from quantities you derive.

Then report the decision tree result in the language of the course and attach the relevant uncertainty, limitation or decision consequence.

Build a representation check before solving. Put classification, nearest-neighbour rule and decision tree into a small symbol-and-units table, mark which values are observed and which are calculated, and predict the direction of the result before doing arithmetic.

A sign, scale or unit mismatch in classification then becomes visible at setup instead of being hidden inside a polished final number.

Run one sensitivity test after the baseline answer. Change the input most closely connected to nearest-neighbour rule, hold the remaining assumptions fixed and recompute only the affected steps. Explain whether the movement in decision tree matches the mechanism.

This nearest-neighbour rule sensitivity shows which assumption controls the conclusion and prevents a single scenario from being presented as universal.

Test with decision tree

Use a three-column classification error log for COSC2670: translation error, calculation error and interpretation error.

Record the exact line where the nearest-neighbour rule solution first diverged, rewrite that line, and check it with a limiting case or an independent calculation.

Correcting the first failed nearest-neighbour rule move is more useful than copying the complete solution again.

A complete response should make the task visible before the detail: identify what must be decided, define the relevant terms, connect the evidence to nearest-neighbour rule, and use decision tree to test the result.

The final sentence about decision tree should answer the question actually asked rather than merely repeat the topic.

The controlling limit is specific: Perfect training accuracy can indicate overfitting and does not establish out-of-sample performance.

Keep that decision tree limit beside the worked example, because it separates a careful COSC2670 answer from one that sounds confident but claims more than the task or evidence supports.

For revision, retrieve classification, nearest-neighbour rule and decision tree without notes, explain their relationship aloud, then complete a changed version of the application: compare a distance-based classifier with a tree while controlling feature scale, depth and validation.

Record the first failed nearest-neighbour rule reasoning move and repair it before attempting another case.

In this chapter

What this chapter covers

  • 01

    classification

  • 02

    nearest-neighbour rule

  • 03

    decision tree

  • 04

    Applying classification

  • 05

    Limits of nearest-neighbour rule and decision tree

Worked example · free

Compare two classifiers under asymmetric cost

Q [5 marks]. A fraud dataset is imbalanced and missing a fraud case costs far more than reviewing a legitimate case. How should k-nearest neighbours and a decision tree be compared?
  • 1Create the same stratified validation folds and scaling pipeline for both models.
  • 1Tune k for the neighbour model and depth or leaf size for the tree inside training folds.
  • 1Evaluate recall, precision and a cost-weighted confusion matrix at candidate thresholds.
  • 1Select the model-threshold pair using the stated error costs, then confirm it once on untouched data.
  • 1Inspect subgroup errors before recommending deployment.
Choose on expected error cost rather than raw accuracy; the final report must include the threshold, untouched-test result and subgroup behaviour.
Sia tip — Scaling is part of the neighbour model, while threshold choice is part of the business decision; validate both without leaking the test set.
Glossary

Key terms

classification
Supervised learning that assigns observations to predefined outcome categories from labelled examples. Use this definition when the task is to compare a distance-based classifier with a tree while controlling feature scale, depth and validation.
nearest-neighbour rule
A classifier that predicts from labels of observations close under a selected distance and feature scale. Use this definition when the task is to compare a distance-based classifier with a tree while controlling feature scale, depth and validation.
decision tree
A sequence of feature-based splits whose branches partition observations toward class predictions. Use this definition when the task is to compare a distance-based classifier with a tree while controlling feature scale, depth and validation.
FAQ

Classification Neighbours and Decision Trees FAQ

What is the main task in Classification Neighbours and Decision Trees?

Compare a distance-based classifier with a tree while controlling feature scale, depth and validation.

How do classification and nearest-neighbour rule work together?

Use classification to establish the object or condition, then use nearest-neighbour rule to explain how it changes the outcome being analysed.

What must a COSC2670 answer qualify here?

Perfect training accuracy can indicate overfitting and does not establish out-of-sample performance.

How should I revise Classification Neighbours and Decision Trees?

Retrieve classification, nearest-neighbour rule and decision tree, apply them to a changed case, and correct the first point where the evidence no longer supports the conclusion.

Study strategy

Assessment move

Reconstruct the relationship among classification, nearest-neighbour rule and decision tree; complete the chapter application without notes; then test the result against this limit: Perfect training accuracy can indicate overfitting and does not establish out-of-sample performance.

Working through Classification Neighbours and Decision Trees in COSC2670? Sia is AskSia’s AI Data Science tutor — ask any COSC2670 Classification Neighbours and Decision Trees question and get a clear, step-by-step explanation grounded in how COSC2670 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.

A+Everything unlocked
Unlocks this Bible + all 6 of your RMIT University subjects - and 1,000+ Bibles across every Australian university.
Sia - your COSC2670 tutor, unlimited, worked the way the exam marks it
The full 2-page Bible + practice bank with worked solutions
Chrome extension - sync your LMS so Sia knows your deadlines
Bilingual EN / Chinese on every Bible and every Sia answer
$0.99 Trial
30-day money-back · cancel in one tap · how it works
COSC2670 · Practical Data Science with Python - independent study guide on the AskSia Library. More RMIT University subjects · Microeconomics across all universities
Unlock the full COSC2670 Bible + 6 RMIT University subjects
$0.99 Trial