By the ZoomCap team.We make ZoomCap, which appears in this article. Prices and features were checked on each vendor's own site on September 29, 2026. Where we could not test a tool ourselves, we say so. How we test
Short answer
ScreenCaptureKit captures system audio when you set SCStreamConfiguration.capturesAudio = true (macOS 13+), and the microphone when you set captureMicrophone = true (macOS 15+). Audio arrives as CMSampleBuffers on your SCStreamOutput with the output type .audio or .microphone, or you attach an SCRecordingOutput (macOS 15+) and let the system write the file. The hard parts are permissions, app-level audio filtering and keeping two audio clocks in sync. If you only need audio from specific apps, Core Audio taps (macOS 14.2+) are the lighter tool.
Where this comes from: the API facts are checked against Apple's documentation, WWDC sessions and release notes, linked throughout and listed at the end. The practical tips come from building the native ScreenCaptureKit recorder in the ZoomCap CLI; we say so wherever a tip is ours rather than Apple's.
TL;DR
- System audio:
capturesAudio,sampleRate(8/16/24/48 kHz),channelCount(1 or 2) andexcludesCurrentProcessAudio, all macOS 13+. - Microphone:
captureMicrophoneandmicrophoneCaptureDeviceID, macOS 15+, delivered as a separate.microphoneoutput. - Audio filtering is per app, never per window. Excluding one window of an app removes all of that app's sound.
- SCRecordingOutput (macOS 15+) writes an MP4 for you. It mixes mic and system audio into one track unless you're on macOS 27 and turn that off.
- Two permissions: Screen & System Audio Recording, plus Microphone. From a CLI, both belong to the terminal app.
- Sync on presentation timestamps. The microphone runs on its own clock, and a dropped audio buffer becomes drift unless you fill the gap.
Who this is for
Developers building a Mac screen recorder, a meeting recorder, a streaming tool or anything else that needs the Mac's sound. If you just want to record your screen with audio, you don't need any of this. Read does Cmd+Shift+5 record audio? instead, since macOS 27 finally added system audio to the built-in toolbar.
The audio APIs, by macOS version
Everything audio-related in ScreenCaptureKit, with the macOS version Apple's documentation lists for each. (Apple's docs now also list iOS 27 and other platforms for much of the framework. This article is about macOS.)
| API | What it does | macOS |
|---|---|---|
capturesAudio | Turns on system audio capture. Off by default. | 13.0 |
sampleRate | 8000, 16000, 24000 or 48000. Anything else falls back to 48 kHz. | 13.0 |
channelCount | 1 (mono) or 2 (stereo). Defaults to stereo. | 13.0 |
excludesCurrentProcessAudio | Drops your own app's sound. Defaults to false. | 13.0 |
SCStreamOutputType.audio | System audio buffers, wrapping an AudioBufferList. | 13.0 |
captureMicrophone | Turns on microphone capture. | 15.0 |
microphoneCaptureDeviceID | Which mic, as an AVCaptureDevice.uniqueID. | 15.0 |
SCStreamOutputType.microphone | Microphone buffers, delivered separately from system audio. | 15.0 |
SCRecordingOutput | Writes the stream (video and audio) to a file for you. | 15.0 |
mixesAudioWithMicrophone | On SCRecordingOutputConfiguration. Set false for two audio tracks. | 27.0 |
If you only need audio and no pixels, ScreenCaptureKit isn't your only option. Core Audio taps (AudioHardwareCreateProcessTap, macOS 14.2+) capture the output of specific processes. They need an NSAudioCaptureUsageDescription string and ask for the narrower audio-only recording permission.
How system audio capture works
Set capturesAudio on the stream configuration and add either a stream output for .audio or an SCRecordingOutput. ScreenCaptureKit then delivers system audio as CMSampleBuffers. Each buffer wraps an AudioBufferList whose format comes from your sampleRate and channelCount. The framework delivers up to 48 kHz stereo.
Audio is filtered by app, not by window
This is the rule that surprises people. In Apple's WWDC22 session, the framework's audio policy "always works at the app level". Capture one Safari window with a single-window filter and the audio includes every Safari tab that's playing. Exclude one Safari window from a display filter and all of Safari's audio disappears. Build your UI around apps, not windows, when audio is on.
Leave your own sounds out
excludesCurrentProcessAudio drops audio from your own process. Turn it on if your recorder plays a countdown beep or a start chime. Leave it off if your app is the thing being demoed.
Capturing the microphone (macOS 15+)
Set captureMicrophone = true, optionally choose a device with microphoneCaptureDeviceID, and add an output for .microphone. Apple's WWDC24 example sets the device to AVCaptureDevice.default(for: .audio)?.uniqueID. Leaving it unset records the system default input, which surprises users who picked a different mic in your app. Pass their choice explicitly.
The microphone is not just another audio stream. An Apple DTS engineer on the developer forums put it this way: with the mic on, app audio and mic audio "arrive with different CMFormatDescriptions", the mic "may arrive at a different rate depending on the input device", and the two are "on independent clocks". Writing both into one AVAssetWriterInput corrupts the file. Give each its own input.
On macOS 13 and 14 there is no captureMicrophone. You capture the mic separately with AVCaptureSession or AVAudioEngine and line it up by timestamp yourself.
SCRecordingOutput or your own AVAssetWriter?
Use SCRecordingOutput unless you need to touch frames or audio samples. It turns a stream into a finished movie file with a configuration object and three delegate callbacks. The default container is MPEG-4 with H.264, and you can choose HEVC or .mov.
| SCRecordingOutput | SCStreamOutput + AVAssetWriter | |
|---|---|---|
| macOS | 15+ | 12.3+ (audio 13+) |
| Code you own | A config and a delegate | Writer, inputs, timing, finalizing |
| Access to samples | No | Every buffer |
| Audio tracks | One mixed track; two on macOS 27 with mixesAudioWithMicrophone = false | Whatever you build |
| Timing bugs | Mostly Apple's problem | Yours |
One tip from our recorder if you use SCRecordingOutput and never read pixels. Capture in kCVPixelFormatType_420YpCbCr8BiPlanarVideoRange rather than BGRA. BGRA is 32 bits per pixel against 12, and it adds a colour conversion before encoding. On a large Retina display at 60 fps, that is the difference between smooth scrolling and dropped frames.
A minimal recorder (illustrative sketch)
This is a sketch to show the shape of the API, not production code. It records the main display with system audio and, if permitted, the microphone, to an MP4. It needs macOS 15 or later.
import AVFoundation
import ScreenCaptureKit
@available(macOS 15.0, *)
final class AudioScreenRecorder: NSObject,
SCStreamDelegate, SCRecordingOutputDelegate {
private var stream: SCStream?
func start(to url: URL, wantsMic: Bool) async throws {
// Requires Screen Recording permission.
let content = try await SCShareableContent.excludingDesktopWindows(
false, onScreenWindowsOnly: true)
guard let display = content.displays.first else { return }
// Audio is filtered per app, so this includes every app's sound.
let filter = SCContentFilter(display: display, excludingWindows: [])
let config = SCStreamConfiguration()
config.width = display.width * 2 // assumes a 2x display
config.height = display.height * 2
config.minimumFrameInterval = CMTime(value: 1, timescale: 60)
config.capturesAudio = true // system audio, macOS 13+
config.sampleRate = 48_000 // 8k, 16k, 24k or 48k
config.channelCount = 2 // 1 or 2
config.excludesCurrentProcessAudio = true // skip our own sounds
// The mic has its own permission: record without it, don't fail.
var useMic = wantsMic
let micStatus = AVCaptureDevice.authorizationStatus(for: .audio)
if useMic && micStatus != .authorized {
useMic = await AVCaptureDevice.requestAccess(for: .audio)
}
config.captureMicrophone = useMic // macOS 15+
// Optional: config.microphoneCaptureDeviceID = device.uniqueID
let stream = SCStream(filter: filter, configuration: config,
delegate: self)
let recordingConfig = SCRecordingOutputConfiguration()
recordingConfig.outputURL = url // defaults: MPEG-4, H.264
let recording = SCRecordingOutput(configuration: recordingConfig,
delegate: self)
try stream.addRecordingOutput(recording)
try await stream.startCapture() // also starts recording
self.stream = stream
}
func stop() async throws {
try await stream?.stopCapture()
// Not done yet: wait for recordingOutputDidFinishRecording.
}
func recordingOutputDidStartRecording(_ output: SCRecordingOutput) {
// output.recordedDuration = video already written.
}
func recordingOutput(_ output: SCRecordingOutput,
didFailWithError error: Error) {
// Surface this. Assume the file is unusable.
}
func recordingOutputDidFinishRecording(_ output: SCRecordingOutput) {
// Now it is safe to open, move or upload the file.
}
func stream(_ stream: SCStream, didStopWithError error: Error) {
// Also fires when the user stops sharing from the system UI
// (SCStreamError.userStopped). Treat that as a normal stop.
}
}If you write the file yourself, the part that matters for audio is the routing. Register each type with addStreamOutput(_:type:sampleHandlerQueue:) and give audio its own queue, so slow video work can't starve it:
// Inside your SCStreamOutput (macOS 15+ because of .microphone).
func stream(_ stream: SCStream,
didOutputSampleBuffer sampleBuffer: CMSampleBuffer,
of type: SCStreamOutputType) {
guard sampleBuffer.isValid else { return }
switch type {
case .screen:
// .idle means the display didn't change; only .complete is new.
guard isCompleteFrame(sampleBuffer) else { return }
startSessionIfNeeded(at: sampleBuffer.presentationTimeStamp)
append(sampleBuffer, to: videoInput)
case .audio:
append(sampleBuffer, to: systemAudioInput)
case .microphone:
append(sampleBuffer, to: micInput) // never the .audio input
@unknown default:
break
}
}
func isCompleteFrame(_ sampleBuffer: CMSampleBuffer) -> Bool {
let attachments = CMSampleBufferGetSampleAttachmentsArray(
sampleBuffer, createIfNecessary: false)
guard let info = (attachments as? [[SCStreamFrameInfo: Any]])?.first,
let raw = info[.status] as? Int,
let status = SCFrameStatus(rawValue: raw) else { return false }
return status == .complete
}Set expectsMediaDataInRealTime = true on each input. The macOS 27 SDK deprecates it in favour of the new input receiver's appendImmediately, but it still works on everything you're likely to support. When an input isn't ready for more data, count the drop. Don't silently discard it (see sync).
Permissions, and the Sequoia prompts
ScreenCaptureKit needs the Screen & System Audio Recording permission, and microphone capture also needs the Microphone permission. They are separate grants in System Settings > Privacy & Security. Apple's user guide notes that people can allow an app to record "both your screen and audio, or just your audio".
- Screen recording: add an
NSScreenCaptureUsageDescriptionstring. Apple's own sample warns that after granting, the app must be restarted before capture works. Tell your users that. - Microphone: add
NSMicrophoneUsageDescriptionand thecom.apple.security.device.audio-inputentitlement. CheckAVCaptureDevice.authorizationStatus(for: .audio)before you configure the stream. - Prefer the system picker. Apple recommends
SCContentSharingPicker(macOS 14+) over building your own selection UI.
The recurring prompt in macOS Sequoia
macOS 15 added a recurring alert for apps that capture the screen directly. It reads "[App] is requesting to bypass the system private window picker and directly access your screen and audio", with an Allow For One Month button. It was weekly in the betas and monthly at release. The macOS 15.1 notes add that people "will see fewer dialogs if they regularly use apps in which they have already acknowledged and accepted the risks."
Apple's macOS 15 notes tie these alerts to deprecated APIs such as CGDisplayStream and CGWindowListCreateImage, and tell developers to move to ScreenCaptureKit and SCContentSharingPicker. Press coverage at the time reported ScreenCaptureKit apps seeing it too. Managed Macs can suppress the alert with the MDM key forceBypassScreenCaptureAlert (macOS 15.1), and VNC-style apps can apply for the Persistent Content Capture entitlement. We couldn't verify from Apple how often the prompt appears on macOS 26 or 27, so design for it. When capture fails to start, point users to the right settings pane instead of showing a raw error.
Capturing from a CLI: the terminal owns the permission
If your recorder is a command-line binary, macOS attributes capture to the app that launched it. That app is the "responsible process": Terminal, iTerm2, or your IDE's terminal. That's where the Screen Recording and Microphone grants go, and it's the entry users need to toggle.
We lean on this on purpose. The ZoomCap CLI ships its recorder as a plain universal binary rather than a .app bundle, because a bundle gets its own privacy identity and needs its own separate grant. A plain binary inherits the terminal's. Two things we learned along the way:
- Don't gate on
CGPreflightScreenCaptureAccess(). For a plain binary it can report the wrong answer for the identity that actually records. We callCGRequestScreenCaptureAccess()as a nudge, then treatstartCapture's error as the source of truth. - Put a timeout on permission probes. On some macOS builds,
getShareableContenthangs rather than failing when the grant is missing. Our check gives up after five seconds, and it runs from the same binary that records, because the answer depends on who asks.
Developers have also reported on Apple's forums that, since macOS Tahoe 26.1, plain executables no longer show up in the Screen Recording list. Point users at their terminal app, not at your binary.
Keeping audio and video in sync
Use each buffer's presentation timestamp, never the time it arrived. Screen, system audio and microphone buffers come in on different queues at different rates, so arrival order tells you nothing. Everything below follows from that.
- Start the writer session at the first complete frame. Apple's docs for
startSession(atSourceTime:)say samples earlier than the start time are written but don't play, and an input whose first sample is later gets an empty edit "to preserve synchronization between tracks". - Treat the mic as its own clock. Apple DTS's advice is to offset mic and app audio relative to your recording start. If the mic's first timestamp is far from the video's, its audio lands before time zero or after the end, and plays as silence. Log the first timestamp of each output type while you develop.
- Never drop audio silently. If you discard a buffer because
isReadyForMoreMediaDatais false, every later sample plays early by that amount. Audio drifts ahead of the picture. Deliver audio on its own queue, and fill any gap between one buffer's end and the next buffer's start with silence. - Expect idle frames. On a static screen, many frames don't carry a
.completestatus. Your audio keeps going while your video track may end early. Repeat the last frame at the end if the track lengths must match.
With SCRecordingOutput, the system handles audio and video interleaving. You still need a clock of your own if you record anything alongside, such as input events or a webcam. We anchor "video time zero" in recordingOutputDidStartRecording: we take the current host time and subtract recordedDuration, rather than trusting when the callback happened to run. We stamp events with NSEvent.timestamp (the same boot-time clock as CACurrentMediaTime()), not the time our handler ran. Under encoding load, handlers run late, and that lateness shows up as drift.
What the ZoomCap recorder does, end to end
For a concrete reference, here is how our CLI's native recorder puts the pieces together:
- Native on macOS 15+.
npx zoomcap recordchecks the OS version first. Older macOS falls back to an ffmpeg recorder that captures video only, becausecaptureMicrophoneandSCRecordingOutputdon't exist there. - One stream, one file. A display filter with no exclusions,
capturesAudiofor system audio,excludesCurrentProcessAudio, andcaptureMicrophonewhen allowed, all written bySCRecordingOutputto an MP4. - The mic never blocks a take. If Microphone isn't granted, we request it for next time, warn, and record screen and system audio without it.
- Every stop is a stop. If the stream ends for a reason we didn't start, we finalize at once. That covers the system Stop control, a revoked permission or an unplugged display. Otherwise the file never closes and the privacy indicator stays lit. A 120-second backstop covers a finalize that truly hangs. Flushing a long, high-res recording can legitimately take tens of seconds, so the backstop can't be shorter.
- Duration comes from the recording. We report
recordedDurationat finish, not wall-clock time, because teardown time isn't video.
Common pitfalls
- No audio output registered.
capturesAudioalone does nothing for a custom pipeline. You also needaddStreamOutput(_:type: .audio, ...), and the same goes for.microphone. - Silence because the app is excluded. Excluding any window of an app removes all of its audio.
- Silence from your own app.
excludesCurrentProcessAudiois on, but the sound you wanted came from your process. - Sample-rate assumptions. Asking for 44.1 kHz gives you 48 kHz. The mic may not match system audio. Read the format description of each buffer.
- A non-exhaustive switch. Apple's own sample handles
.screenand.audio. Once you turn the mic on, handle.microphoneand add@unknown default. - Exiting before the file is finished. Wait for
recordingOutputDidFinishRecording(orfinishWriting) before you quit, move or upload. - Granting the wrong app. From a CLI, the terminal needs the grant, and it may need a restart after.
- Queue depth over 8. Apple says not to exceed eight frames. We use eight on large Retina displays so a slow encoder doesn't drop frames.
When ScreenCaptureKit is the wrong tool
- Audio only, from specific apps: use Core Audio taps. There are no pixels to pay for, and the permission is narrower.
- You just need a recording, not a product: on macOS 15 or later,
npx zoomcap recordcaptures screen, system audio and mic in one command; on macOS 27, Cmd+Shift+5 records system audio too, and OBS is the free option on older systems. Our guide to recording a Mac screen with audio covers the rest. - A still bug report: a screenshot and three lines of text beat a video nobody wants to scrub through.
Our recommendation
Target macOS 15 or later if you can. You get the mic in the same stream and SCRecordingOutput for the file, which removes most of the sync code you'd otherwise write. Drop to SCStreamOutput and AVAssetWriter only when you need samples or separate tracks before macOS 27. Design your permission flow for a terminal as carefully as for an app, because that's where CLI users get stuck.
If you'd rather not build a recorder at all, ours is one command. npx zoomcap record captures the screen, system audio and your mic on macOS 15+, and the browser editor adds zooms where you clicked. It's part of the paid plans, $29 once or $12.49 a month. The recording is edited and exported on your machine with nothing uploaded; the security page shows the data flow. It records the whole display with no per-app audio choice, and on Linux it records video only. The docs list every flag. If you need scenes, streaming or per-app audio, OBS is free and does them; we compared the two in OBS vs ZoomCap.
Frequently asked questions
Which macOS version does ScreenCaptureKit need to capture the microphone?
macOS 15 Sequoia. captureMicrophone, microphoneCaptureDeviceID and the SCStreamOutputType.microphone output were all introduced in macOS 15.0. System audio capture (capturesAudio) has been available since macOS 13 Ventura.
Can ScreenCaptureKit capture audio from a single app?
Yes, but only at the app level. Audio follows the content filter per application: include only that app's windows and you get only its sound. You can't isolate one window's audio, because a single-window filter captures all audio from the app that owns the window. For audio-only capture of specific processes, Core Audio taps (macOS 14.2+) are the more direct tool.
Does SCRecordingOutput put system audio and the mic on separate tracks?
By default it mixes them into one audio track. macOS 27 added SCRecordingOutputConfiguration.mixesAudioWithMicrophone; set it to false to keep system audio and microphone as two tracks. On macOS 15 and 26 there is no such switch, so if you need separate tracks there, write the file yourself with AVAssetWriter.
Why is my ScreenCaptureKit audio track silent?
The usual causes: capturesAudio is false; you never added an output for .audio; the app making the sound is excluded by your content filter; the sound comes from your own process and excludesCurrentProcessAudio is true; or the buffers were written with timestamps outside your AVAssetWriter session. Log the first presentation timestamp of each output type and compare.
Does ScreenCaptureKit work from a command-line tool?
Yes. A plain binary launched from a terminal is treated as part of that terminal for privacy purposes, so the Screen Recording and Microphone grants belong to Terminal, iTerm2 or whichever app launched it. Tell users to grant the terminal, and restart it if capture still fails after granting.
Sources: Apple's ScreenCaptureKit documentation, the Capturing screen content in macOS sample, WWDC22 sessions 10156 and 10155, WWDC24 10088, and the macOS 15 and 15.1 release notes. All were checked in September 2026.
Skip the build and record today. Try ZoomCap free in the browser recorder, or get the ScreenCaptureKit CLI with system audio, mic and automatic zoom for $29 once on the pricing page.