The HTML kind attribute
The HTML kind attribute sets what a <track> element provides, one of subtitles, captions, descriptions, chapters or metadata. A missing value means subtitles, and an unknown value means metadata, which the browser never displays.
Overview
The kind attribute declares how a timed text track is meant to be used. It sits on the <track> element inside <video> or <audio> and decides whether the browser presents the cues to viewers or leaves them to your scripts.
Three values are for viewers. subtitles carry a transcription or translation of the dialogue for people who can hear the audio but not understand it. captions add sound effects, music cues and other audio information for people who cannot hear it, and descriptions describe the picture and are meant to be spoken aloud.
The other two are for code. The current HTML Standard describes chapters and metadata tracks as intended for use from script and not displayed by the browser. MDN's compatibility data also shows no browser that implements descriptions tracks yet, so only captions and subtitles appear on screen.
The defaults are easy to trip over. Leave kind out and the track counts as subtitles, which makes srclang required. Misspell it, as in kind="caption", and the invalid-value default turns it into metadata, so the text quietly never shows.
Syntax
<track kind="captions" src="captions-en.vtt" srclang="en" label="English (CC)">
Values
| Value |
|---|
| subtitles | captions | descriptions | chapters | metadata |
Example
<video id="v" src="https://codeshack.io/web/example.mp4" controls muted width="240">
<track kind="captions" srclang="en" label="English (CC)" default src="data:text/vtt,WEBVTT%0A%0A00:00.000%20--%3E%2000:31.000%0A(soft%20music)%20Captions%20describe%20sounds%20too">
<track id="meta" kind="metadata" src="data:text/vtt,WEBVTT%0A%0A00:00.000%20--%3E%2000:31.000%0Ascene=intro">
</video>
<p id="out">Metadata cue: none yet</p>
<p><small>Each track inlines its WebVTT text as a data: URL. A real page points src at a .vtt file.</small></p>
<script>
const meta = document.getElementById('meta').track;
meta.mode = 'hidden';
meta.addEventListener('cuechange', () => {
const cue = meta.activeCues[0];
document.getElementById('out').textContent = 'Metadata cue: ' + (cue ? cue.text : 'none');
});
</script>
Best practices
- Use
captions, notsubtitles, for same-language text aimed at deaf and hard-of-hearing viewers, and include sound effects and speaker changes in it. - Always write
kindout, since a missing value falls back tosubtitlesand a typo falls back to the hiddenmetadatakind. - Add srclang to every subtitles track, because the HTML Standard requires it for that kind.
- Do not expect
chaptersordescriptionstracks to appear on their own. Browsers do not render them, so build any chapter menu from the cues with script. - Use
metadatatracks for timed data your code needs, such as scene markers, and read the cues from the cuechange event.
Accessibility
The kind value decides who a track serves. W3C technique H95 uses a captions track to meet WCAG success criterion 1.2.2 for prerecorded video, and it notes that a subtitles track is not enough when sound effects or other audio matter.
Descriptions tracks are meant to be synthesized as audio, for example for blind viewers, but no browser implements them yet according to MDN's compatibility data. Do not count on a descriptions track to provide audio description for a video.