Press play. Every word lights up as it's spoken, with every speaker identified. This is actual public-record body-worn camera footage, transcribed exactly the way your discovery is.
The transcript above is the deliverable: a clean, court-ready Word document in exactly that format: speaker-attributed, timestamped to the recording and the real-time clock from the camera.
Your transcript will not look like this sample unless we receive all of the relevant discovery for the case. Speaker names, correct terminology, and the real-time clock [in brackets] are reconstructed from the materials you provide:
Without complete discovery, speakers can only be labeled generically (e.g. "Unknown Male") and proper nouns may be approximate. The more discovery you provide, the more precise the names and speaker labels: every officer and every subject, identified.