The University of Hong Kong · FACULTY OF MANAGEMENT

PMGM7023 Chap.6 Evidence, Tools and Control in AI Systems

- one subject, every graph, every model, every mark
4 Chapters3-page Bible
Our own words - no uploaded lecturer files
Updated for this semester
Chapter 6 of 12 · PMGM7023

Evidence, Tools and Control in AI Systems

Fluency is not support

The third layer of the vocabulary exists because the most expensive failures in this course are not obviously wrong. They are fluent, plausible and specific: an invented reference complete with a journal and an identifier, a filled-in eligibility rule the source never contained, a trend described confidently that the supplied table does not contain.

The mechanism is ordinary.

The system is asked to continue an answer, the needed evidence is missing, weak or contradictory, a statistically plausible continuation is selected, and fluent wording hides that nothing was verified. A larger or more articulate system is not automatically more truthful, since it can also reproduce false patterns present in the text it learned from.

Nothing in that chain involves intent, which is why reading for confidence cannot detect it. The fast test is to ask it to point at the line that carries the support: a system that can point at the line has read something, and one that restates the claim in different words has not.

Bringing outside evidence into the answer

Retrieval-augmented generation is the standard repair.

A question-time process searches a collection the application is permitted to use, adds the relevant passages into the current context, and asks the system to answer from those passages and cite them. The sources stay outside the trained model, which lets an organisation update a policy without retraining anything.

The limitation matters as much as the mechanism: if the search returns nothing relevant, or returns a superseded version, the answer is grounded in the wrong thing and still arrives with a citation attached.

Two ways a task reaches other software

An application programming interface is one service's own request-and-response contract: you state the address to call, the operation you want and the data you hand over, and a credential identifies the calling application so the service can check permission and record usage.

A shared protocol is a common format through which an AI application discovers and uses tools offered by a server, so one client can work with many servers in the same message grammar. The first is one service's contract; the second is a grammar for many. Neither is spoken natively by the underlying systems, since a server sits in between and exposes only what it is configured to provide.

Two further terms belong here.

A reusable bundle packages instructions, steps, examples, templates and checks for a recurring kind of work, storing know-how rather than performing an operation. An agent is a system that selects and uses tools across several steps toward an objective, which a single-answer assistant is not.

An enforceable limit, not a polite request

The fifth layer separates the surrounding machinery from the model.

A harness is the software and instructions that organise how work is performed and controlled: what information reaches the model, the steps and handoffs, which tools are available, which permissions and confirmations apply, and the logging, retries and stopping conditions. A guardrail is a rule, filter, permission, validation or halt that is actually enforced, and the distinction is testable.

An appeal to take care is only an instruction, while a tool call that will not run, or an approval the work waits on, is a guardrail. Checks can sit before the input, during the action or after the output, and only the middle position can stop something before it happens.

In this chapter

What this chapter covers

  • 01

    Unsupported fluency, its mechanism and the test that detects it

  • 02

    Retrieval at question time, and what it still cannot fix

  • 03

    A service contract against a shared tool protocol

  • 04

    Harness, guardrail, and the difference between advice and enforcement

Worked example · free

Diagnosing an invented policy rule

Q [8 marks]. AskSia-authored practice. A service team's assistant tells customers that refunds above a stated amount require manager approval. No such rule exists in the company policy. Name the layer, the repair, and one repair that would not work. The marks shown are an AskSia study allocation, not an official University marking scheme.
  • 3Attribute the failure to a layer and justify the attribution.
  • 3State the repair inside that layer.
  • 2Name a plausible repair that would not address it.
This is an evidence-layer failure. The question required a specific eligibility rule, the policy supplied did not contain one, and a plausible threshold completed the answer. The repair is to search the current policy at question time, pass the relevant clauses into the context, require every rule stated to quote the clause it came from, and permit the answer to say the policy is silent. A check comparing each stated rule against the passages used catches the remainder. Adjusting the sampling setting would not help, because the failure is missing evidence rather than variable wording.
Sia tip — Before proposing a fix, say which layer produced the failure. A repair applied to the wrong layer looks like diligence and changes nothing.
Glossary

Key terms

Application Programming Interface
An application programming interface is one service's published contract for receiving a request and returning data. Each service defines its own.
API Key
An API key is a secret credential sent with a request so a service can identify the caller, check its permission and record usage against an account.
Harness
A harness is the software and instructions surrounding a model that organise how work is performed, which tools are reachable and when execution stops.
Recursive Self-Improvement
Recursive self-improvement is an arrangement in which a system helps improve its own capabilities and the improved version contributes to the next round. It describes a feedback mechanism, not a guarantee of progress.
FAQ

Evidence, Tools and Control in AI Systems FAQ

How can I tell whether an AI answer is actually supported?

Ask it to point at the line that carries the support. A system that has read or searched something can point at the passage; one that has produced a plausible continuation will restate the claim in different words. Polished wording is not proof, and a specific-looking reference with a journal and an identifier is one of the most common shapes an unsupported claim takes.

What makes something a guardrail rather than an instruction?

Enforcement. A blocked tool call, a restricted folder, a required human approval before a high-impact action or a validation that rejects output are guardrails because something prevents the behaviour. A sentence asking a system to be careful, to avoid errors or to stay within scope is an instruction, and nothing stops it being ignored. Checks can sit before input, during action or after output.

Study strategy

Exam move

Build a one-page map of the five layers with one failure example under each, and rehearse attributing described failures to the right layer. When you meet an unfamiliar term, write the system question it answers rather than a definition. Practise the evidence test on any confident claim you encounter, including your own drafts.

Working through Evidence, Tools and Control in AI Systems in PMGM7023? Sia is AskSia’s AI Management tutor — ask any PMGM7023 Evidence, Tools and Control in AI Systems question and get a clear, step-by-step explanation grounded in how PMGM7023 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.

A+Everything unlocked
Unlocks this Bible + all 4 of your The University of Hong Kong subjects - and 1,000+ Bibles across every Australian university.
Sia - your PMGM7023 tutor, unlimited, worked the way the exam marks it
The full 3-page Bible + practice bank with worked solutions
Chrome extension - sync your LMS so Sia knows your deadlines
Bilingual EN / Chinese on every Bible and every Sia answer
$0.99 Trial
30-day money-back · cancel in one tap · how it works