CHIN2600 Chap.5 Transcription: Turning Talk into Data
Transcription: Turning Talk into Data
Transcription looks mechanical and is not. Every convention you switch on is a decision that something deserves a reader's attention, and every one you leave off is a decision that it does not.
The system taught in this course makes that decision explicit by distinguishing two levels.
A broad transcript carries the words, the speakers, the division into turns and intonation units, truncation of both, the contours, laughter, pauses of medium and long duration, and any stretch you could not hear clearly. A narrow transcript adds accent, tone, prosodic lengthening, breathing and other vocal noises.
The working rule is broad across the whole extract and narrow only on the stretches your analysis rests on.
The working subset is small: one intonation unit per line with a speaker label, transitional continuity at each unit boundary, aligned brackets for overlap, timed pauses where the claim needs them, the transcriber's own marks for comments, uncertain hearings and indecipherable syllables, paired tags for prosodic quality, and a tagged span for code switching, which in this course is often the phenomenon itself.
Two notations coexist here and they are not interchangeable: the discourse transcription set used for analysis, and a corpus format with its own labels and timestamps used for the recorded telephone exercise.
Mixing symbols from the two inside one transcript is the easiest way to lose credibility on a slide. Chinese extracts are laid out in four lines: characters, romanisation with tone numerals, a word by word gloss, and an idiomatic translation.
What this chapter covers
- 01
5.1 Broad against narrow, and deciding which stretch gets which
- 02
5.2 The seven step build order for a usable transcript
- 03
5.3 The working subset of symbols, and when each is actually needed
- 04
5.4 The four line layout every Chinese extract needs
- 05
5.5 Two notations in one course, and why they must not be mixed
Choosing the level of detail for one analytic claim
- +2Turn boundaries, because the claim is entirely about what follows what.
- +3The pause after the repeat, timed in seconds, because a short gap and a three second gap support opposite readings.
- +3Transitional continuity at the end of the child's repeat, since a final contour invites closure while an appealing one makes a response due and therefore makes its absence meaningful.
- +2Nothing else. Marking voice quality across the whole extract would add pages and change nothing, and heavy notation on unused material reads as decoration.
Key terms
- Intonation unit
- A stretch of speech produced under a single coherent intonation contour, written on its own line. Unit boundaries show how a speaker packages what they are saying, which is why the layout does analytic work before any symbol does.
- Transitional continuity
- The degree of continuity marked at the point where one intonation unit passes into the next. Without it a reader cannot hear a list as a list or a question as a question.
- Truncation
- A break before completion, marked differently for a unit whose projected contour is abandoned and for a word whose ending is not uttered. The two are different marks and different findings.
- Broad transcription
- The level that carries words, speakers, turns, units, truncation, contours, laughter, pauses of medium and long duration and stretches heard uncertainly. It is the default for a whole extract.
- Narrow transcription
- The level that adds accent, tone, prosodic lengthening, breathing and other vocal noises. It is applied to the stretches an argument depends on rather than to everything.
- Interlinear gloss
- The four line layout used for Chinese data: characters, romanisation with tone numerals, a word by word gloss with grammatical labels, and an idiomatic translation. It is what lets a reader follow an argument about a single morpheme.
Transcription: Turning Talk into Data FAQ
Does the final video really require transcripts?
Your data analysis has to include transcripts for interactional data or texts such as screenshots for written data, with characters shown for the Chinese material and words for the English. A recording described but not transcribed cannot be analysed on a slide.
Which transcription system does this course actually teach?
The discourse transcription set introduced in the transcription week, with its units, transitional continuity, overlap brackets, timed pauses and paired prosodic quality tags. The separate corpus exercise uses a different format with its own speaker labels and timestamps, and the two should never be mixed in one transcript.
How much detail is too much?
Any detail your caption does not need. Uniformly narrow notation tells a reader you had not decided what your claim was when you made the transcript, and it costs you space that the analysis needs.
What do I do about a stretch I genuinely cannot hear?
Mark it. There are separate conventions for an uncertain hearing and for an indecipherable syllable, and using them is part of the method. Silently repairing a stretch you could not hear is the one move here that is not recoverable.
Assessment move
Transcribe one minute of anything this week, before you have data of your own, and time yourself. Most students discover that a minute takes far longer than they planned, which is the single most useful thing to learn early, because the video deadline does not move. Then write the caption for that minute and delete every mark the caption does not need; what survives is the working subset you will actually use.
Working through Transcription: Turning Talk into Data in CHIN2600? Sia is AskSia’s AI Arts and Humanities tutor — ask any CHIN2600 Transcription: Turning Talk into Data question and get a clear, step-by-step explanation grounded in how CHIN2600 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.
Digital Marketing and Social Media · Organisational Behaviour · Identifying and Assisting Students at Risk