Apple announced the Apple Watch Series 12 and Apple Watch Ultra 4 on September 9, starting at $399 and $799. Both ship on September 18. Both carry a new class of features Apple calls Audio Intelligence, and two of them listen to the conversations happening around you: Live Rewind reads back the last 15 seconds of speech, and Siri Recap takes ambient notes and turns them into summaries.
The bigger story for anyone who works from conversations is that Apple has now entered this category. Ambient AI note taking is no longer a niche accessory idea. Apple put it on the wrist of a product it sells by the tens of millions.
Read Apple's own description closely, though, and the design goal is specific. These features exist to help you stay present in the moment. They do not exist to leave you a record you can use later. That single choice decides where the watch fits and where it runs out.
Three steps sit on that spectrum, and it is worth knowing which one you need. The watch helps you recall a conversation. A dedicated note taker leaves you a record of it. An agent takes that record and does the next thing with it.
The news in brief
- Announced September 9. On sale September 18. watchOS 27 arrives September 14.
- Apple Watch Series 12 starts at $399. Apple Watch Ultra 4 starts at $799.
- Four Audio Intelligence features: Sound Recognition, Shazam music recognition, Live Rewind, and Siri Recap.
- Hardware requirement: the new S11 chip, so Series 12 and Ultra 4 only.
- Live Rewind and Siri Recap arrive in beta later in 2026, in English first, and need an iPhone 16 or later (the iPhone 16e is excluded). Neither is available in the EU at launch.
What Audio Intelligence does on your wrist
Audio Intelligence is a set of four microphone-driven features that run on the S11 chip, with support from Apple Intelligence models on your iPhone and Apple's Private Cloud Compute.
Sound Recognition listens for sirens, alarms, doorbells, and a baby crying, then alerts you. It works even when your iPhone is not nearby, and it brings a real accessibility win to the watch for people who are deaf or hard of hearing.
Music recognition with Shazam identifies songs playing near you and shows the title and artist in a Smart Stack widget.
Live Rewind shows the last 15 seconds of speech as a text snippet when you double-press the Digital Crown. You can ask Siri about the snippet or save it to the Siri app.
Siri Recap takes ambient notes through the day and produces an Apple Intelligence summary with a title and key points, which you review in the new Siri app.
The privacy engineering behind this deserves credit, and it matters for the comparison later. According to Apple's support documentation, raw microphone audio flows into a buffer inside a hardware-isolated Secure Exclave on the S11, where new audio continuously overwrites old audio. No conventional recording file gets created or stored. Every feature is opt-in. Live Rewind plays an audible chime and shows a microphone indicator when you trigger it, even if the watch is silenced. Apple also published an 11-page privacy paper on how all of this works.
That is a serious piece of work, and it is the reason the interesting question is not whether Apple is trustworthy here. The interesting question is what you end up holding when the conversation is over.
Why the coverage split in a day
Coverage of the announcement divided almost immediately, with some outlets framing Audio Intelligence as surveillance hardware and others leading with Apple's on-device processing and opt-in controls. Both sides were reading the same Apple materials.
Engadget and TechCrunch leaned into the always-listening framing. Wired and 9to5Mac led with the privacy architecture. Neither reading is wrong, because the tension is real and sits in the product itself.
Apple's protections are strong for the person wearing the watch. You choose whether the features are on, you choose when, and your unsaved Siri Recap summaries disappear after seven days. The people sitting across from you never make any of those choices. They get a chime and an on-screen indicator when Live Rewind fires, and nothing at all to opt into for Siri Recap. Apple clearly anticipated that reaction, which is why an 11-page privacy paper shipped alongside a smartwatch.
So the honest summary is that the watch handles your data carefully and handles the room's expectations less carefully. Which brings us to the part the reviews have mostly skipped: what these features can and cannot produce.
Where the 15-second limit stops you
Live Rewind gives you the last 15 seconds of speech as text, on demand, and nothing more. It does not transcribe in the background, and it keeps nothing unless you save the snippet yourself.
Picture the moment it is built for. A client says a number. You nod and keep listening. Thirty seconds later you want to check that number. It is gone. Live Rewind requires you to notice the gap within 15 seconds and reach for the Digital Crown while you are mid-conversation.
Three limits follow from the design:
- No background transcription, so you cannot go back to a stretch of conversation that has already passed.
- Nothing persists unless you ask Siri about the snippet or save it to the Siri app.
- You need a free hand and a moment of divided attention, at the exact point you were trying to pay attention.
None of that makes Live Rewind a bad feature. It makes it a memory aid with a 15-second window.
What a Siri Recap summary leaves out
Siri Recap produces a title, a summary, and key points. Apple is explicit that it does not produce a transcript, and three consequences follow that most reviews have not worked through.
It does not tell you who said what. Apple states that Audio Intelligence does not attribute speech to speakers. A meeting summary without attribution cannot give you an owner for any commitment. "Someone agreed to send the revised scope by Friday" is not a follow-up you can act on.
It leaves out the details you most wanted kept. Siri Recap is designed to omit sensitive information such as financial data and personal identification numbers. Read that as a feature, because it is one. Then read it as a limit: the quoted price, the dosage, the account number, and the contract date are exactly the fragments a summary is built to skip.
It clears itself. Unsaved summaries auto-delete after seven days. For a passing conversation that is sensible hygiene. For a clinical visit, an interview, or a negotiation, a seven-day expiry with no transcript and no export is not a record at all.
What it takes to turn these on
Live Rewind and Siri Recap carry a longer requirements list than the keynote suggested, and the gating happens on both the watch and the phone.
- Apple Watch Series 12 or Apple Watch Ultra 4, because both features depend on the S11 chip and its Secure Exclave.
- An iPhone 16 or later, with the iPhone 16e excluded, plus Apple Intelligence enabled.
- Beta availability later in 2026, in English first, with more languages to follow.
- No availability in the EU at launch.
Sound Recognition and Shazam recognition arrive with the watches. The two conversation features arrive later, in beta, for a narrower set of people. If you own a Series 11 and an iPhone 15, none of this reaches you.
Three things Apple has not said
Apple's announcement left real gaps, and each one changes the decision for a different group of people.
- Whether older watches get anything. Apple has not said if Series 11 or earlier models receive any Audio Intelligence feature through watchOS 27. Current owners cannot yet tell whether an upgrade is required.
- How much conversation Siri Recap covers. Apple has not published a length limit, where summaries live, or whether they sync to iPhone. Without that, nobody can judge whether the output works for anything beyond personal recall.
- When the EU gets it. Apple says only that the features are unavailable there at launch. European buyers have no timeline.
We will update this article as Apple fills those gaps.
Compare the two side by side
The clearest way to see the difference is to line up the two approaches on the things you would need after a conversation ends. Plaud NotePin S is a $179 wearable AI note taker, so the price comparison is not watch-to-recorder but capability-to-capability.
Apple Watch Audio Intelligence |
Plaud NotePin S |
|
|---|---|---|
How far back you can go |
15-second text snippet |
Up to 20 hours of continuous recording |
Transcript |
None, summary only |
Full transcript |
Who said what |
Not attributed |
Speaker labels |
Languages |
English first, more to follow |
112 languages through Plaud Intelligence |
Sensitive details |
Summary omits financial and ID data |
Kept, and yours to redact or share |
Retention |
Unsaved summaries auto-delete after 7 days |
Stored, searchable, 64 GB on device |
Export |
Not offered |
Export and sync across Plaud App, Plaud Web, and Plaud Desktop |
Device requirements |
Series 12 or Ultra 4, plus iPhone 16 or later |
Standalone device, no phone model requirement |
Regional availability |
Not in the EU at launch |
Available globally |
Compliance posture |
Not positioned for regulated recordkeeping |
ISO 27001, ISO 27701, GDPR, SOC 2 Type II, HIPAA, EN 18031 |
Price |
From $399 (watch) |
$179 |
Now the fair half of that table. For a large slice of everyday life, the watch is the better answer, and it is better for a reason no accessory can beat: it is already on your wrist. Missing one line in a hallway chat, wanting a song title, needing an alert for a siren you did not hear, all of that suits a wrist feature and suits it well. You do not have to remember to bring anything, charge anything, or wear anything new.
The split shows up when the conversation has consequences.
What a wrist summary cannot give you
A summary reminds you. A record lets you work. Below is what the second one buys you, grouped by what you can do with it.
The whole conversation, not a memory prompt
Plaud NotePin S records up to 20 hours in one session, with 64 GB on board and cloud storage that does not cap out. A three-hour hearing, a full day at a conference, or a 90-minute lecture stays intact from start to finish.
You get a full transcript, not only a summary. That difference decides whether you can quote someone accurately later, check a claim against what was said, or hand a colleague the passage that settled an argument. Everything then exports and syncs across Plaud App, Plaud Web, and Plaud Desktop, so the record lands in the tools you already work in rather than expiring in a phone app.
Who said what, and what they committed to
Speaker labels are the single hardest difference in this comparison, because Apple has ruled attribution out by design. Plaud Intelligence identifies each voice in the room and ties every line to a speaker. Custom vocabulary, currently in beta, covers verticals including education, sales, healthcare, finance, retail, and real estate, which keeps names, drug names, product codes, and contract terms from being transcribed into something else.
Here is an illustrative example of what that produces, not a real client meeting:
Speaker 1 (Sales): We can hold the current rate through the end of Q4.
Speaker 2 (Client): Only if onboarding starts before November 15.
Speaker 1 (Sales): I'll confirm the delivery timeline by Friday.
Action items
- Speaker 1: confirm delivery timeline by Friday
- Speaker 2: decision depends on onboarding starting before November 15
Key data
- Current rate held through end of Q4
Every line in that output rests on three things Apple has chosen not to provide: a transcript, speaker attribution, and the numbers a summary would filter out.
Work you can keep doing after the conversation
The record is the starting point rather than the deliverable. Plaud Intelligence turns one recording into different documents depending on the job: pick from more than 10,000 summary templates, or build your own. The same client visit can produce a meeting summary for your team, a structured clinical note, or a sales follow-up, without listening to the audio again.
Ask Plaud, in beta, lets you question your own archive rather than scroll it, so "what delivery date did we agree with this client last time" becomes a question with an answer. AutoFlow pushes outputs into your workflow, multidimensional summaries produce mind maps, key data, and to-do lists, and multimodal input lets you attach the whiteboard photo or the slide that the audio alone would miss. Audio import brings in recordings made elsewhere.
All of it stays searchable. An archive gets more useful the longer you keep it, which is the opposite of a seven-day expiry.
Reach, in languages and places and phones
Plaud Intelligence transcribes and summarizes in 112 languages. Apple starts with English and adds more later. Plaud works globally, including the EU, where Audio Intelligence is not available at launch.
A Plaud device also stands on its own. No iPhone 16 requirement, no minimum watch generation, nothing to enable in Apple Intelligence. Someone on an iPhone 15 or an Android phone can use it today.
Then there is a category the watch does not enter at all. Audio Intelligence handles the room you are standing in. It does not record phone calls, and it does not capture online meetings. Across the wider Plaud lineup, Plaud Note and Plaud Note Pro handle phone call recording, and Plaud Desktop records online meetings without sending a bot into the call. Plaud NotePin S is built for in-person conversation, so it deliberately leaves calls to those other options.
What the next step looks like with Plaud One
Apple's version of ambient AI stops at recall. Plaud One, an earbud product currently shipping as a limited Explorer Edition, shows what the same idea looks like when it carries on into the work itself.
Read the two side by side and the shape of the gap is clear:
Apple Watch Audio Intelligence |
Plaud One |
|
|---|---|---|
How you start it |
Double-press the Digital Crown |
Long-press the record button, or say "Hey Plaud" |
Asking it something |
Ask Siri about the snippet |
Ask anything, with the answer in your ear or through the case |
What it covers |
The room you are in |
In-person through the case, plus phone calls and online meetings through the earbuds |
Pickup range |
Wrist microphones |
Earbuds up to 6.6 feet, case up to 16.4 feet with 4 MEMS and AI beamforming |
Phone dependency |
iPhone 16 or later required |
Built-in 4G LTE eSIM in the case, so it works with no phone nearby |
What you get out |
A 15-second snippet or a summary |
Transcripts, summaries, and artifacts such as reports, decks, spreadsheets, and web pages |
Acting on it |
Nothing beyond the summary |
Connects to Gmail, Google Calendar, Slack, Notion, Google Drive, and MCP-compatible services |
The last row is the whole difference. Siri Recap finishes when you have remembered something. Plaud One is built to keep going: drafting the follow-up, booking the meeting, writing the document, running a scheduled routine on its own.
So the honest picture is three steps rather than two. The watch helps you recall. A dedicated note taker leaves you a record. An agent takes the record and does the next thing with it.
Be clear about the status before you get attached to it. Plaud One Limited Explorer Edition is $249.99, capped at 2,000 units worldwide, and sold only in the United States, France, Germany, the United Kingdom, Italy, Spain, Canada, and the Netherlands. The current pre-order round has closed and a waitlist is open for the next one. The agent capabilities run on Plaud Intelligence beta features, some of which arrive through later firmware updates, and availability varies by language, region, device, and software version. Treat it as a look at where the category is heading rather than something to buy this week.
If you want the same underlying record today, with a normal purchase path, that is what the rest of this comparison covers.
Who is responsible for the record afterward
The two products place the responsibility for a record in different places, because only one of them sets out to create a record at all.
Give Apple its due first. Raw audio is processed inside a hardware-isolated Secure Exclave on the S11 and discarded, no audio file is created, every feature is opt-in, and Live Rewind announces itself out loud. On the narrow question of whether your voice leaves the device, Apple's approach is strong, and no cloud service can match a design that throws the audio away.
The trade-off arrives afterwards. Apple provides no transcript, no export, and no retention control, and unsaved summaries clear after seven days. For a clinician, a lawyer, an HR lead, or an auditor, that is not a privacy benefit. Those roles are obliged to hold a record, produce it on request, and show how it was handled, and Apple has not positioned Audio Intelligence for regulated recordkeeping.
Plaud is built for the other side of that line, which means accepting the obligations that come with holding a record. The current commitments:
- Security management aligned to ISO/IEC 27001.
- Privacy management aligned to ISO/IEC 27701.
- Built with GDPR-aligned privacy protections.
- Independently audited controls with SOC 2 Type II.
- Healthcare-grade safeguards aligned with HIPAA expectations.
- Built to meet EN 18031 cybersecurity requirements, which covers cybersecurity for wireless hardware.
- Your data is never used for AI training unless you explicitly opt-in.
Full detail sits in the Plaud Trust Center. The point is not that one company is safer than the other. The point is that they are answering different questions, and you should pick based on which question is yours.
There is a second obligation that sits with you rather than with either company. Before you record, take a moment to let others know and get their okay. Most US states require the consent of one party to a conversation and a smaller group requires consent from everyone involved, which the Justia 50-state survey on recording conversations sets out state by state. In the UK and the EU, the ICO's guidance on consent under the UK GDPR covers what counts as valid consent. One sentence at the start of a meeting handles it, and Apple's own decision to give Live Rewind an audible chime says the same thing about where the line sits.
Match the tool to what you need afterward
Ask what you need to exist once the conversation ends, then choose. If you need a nudge, the watch covers it. If you need something you can quote, assign, or file, the watch cannot produce that by design.
An everyday moment you missed
You lost one sentence in a corridor conversation, or you want the name of the song in the café. Use the watch. It is already on you, it needs no preparation, and it asks nothing of you afterwards. A dedicated recorder would be overkill.
Team meetings and project decisions
You need action items with owners. Speaker labels are the deciding factor, and a summary that cannot attribute a commitment leaves you rebuilding the meeting from memory.
Client conversations with numbers in them
Pricing, terms, delivery dates, scope changes. A summary designed to omit financial detail removes the part you needed. A transcript keeps the number and the sentence around it.
Doctor visits, from the patient's side
Dosages, follow-up timing, what to watch for. This case gets overlooked, and it is one of the strongest. You are not going to remember the instructions, and a summary that erases itself in seven days is no help at the follow-up appointment six weeks later.
Interviews, research, and classes
Journalism, user research, and study notes all depend on quoting accurately. Only a transcript supports that, so the watch is out before you start.
Field work and conversations on the move
Site visits, walk-and-talks, and anything where your hands are busy. Plaud NotePin S weighs 0.61 ounces and comes with a lanyard, wristband, clip, and magnetic pin in the box, so you can wear it in whatever way suits the day. One press starts recording, and a second press marks a highlight without stopping the conversation.
Phone calls and online meetings
Neither is covered by Audio Intelligence. If most of your conversations happen through a phone or a video call, look at Plaud Note and Plaud Note Pro for call recording, or Plaud Desktop for online meetings, rather than a wearable.
Four objections, answered
The same worries come up every time someone weighs a dedicated recorder against a feature they already own, and each one deserves a straight answer.
I already wear a watch, so why carry another device? If you need to catch the occasional missed sentence, do not buy anything. Get the Series 12 feature when it ships and you are done. A separate device earns its place when your output depends on the detail inside conversations, which is a different problem from forgetfulness. On the practical worry: 40 days of standby means charging is a weekly thought rather than a nightly one, and 0.61 ounces on a lapel is easy to forget you have on.
Does it cost a subscription? Yes, above a free tier, and it is worth being plain about it since this article has been counting Apple's requirements. Plaud NotePin S is $179 and includes the Plaud Starter Plan free, with 300 minutes of transcription per month. Plaud Pro Plan is $99.99 per year for 1,200 minutes per month, and Plaud Unlimited Plan is $239.99 per year. What matters is that the tiers differ on volume, not on capability. Speaker labels, export, the full template library, and every compliance commitment listed above are all in the free tier. You can see the current plan breakdown on the Plaud Intelligence pricing page.
Apple processes on device and Plaud uses the cloud, so isn't Apple more private? On that specific question, yes. Apple discards the audio inside the chip, and nothing beats not keeping it. The reason Plaud works differently is the whole subject of this article: a full transcript, speaker attribution, a searchable archive, and sync across your devices all require processing beyond the device. That trade comes with commitments. Data is protected in transit with TLS encryption. Your data is encrypted at rest with AWS server-side encryption. Your data is never used for AI training unless you explicitly opt-in.
How accurate are the AI summaries? Accurate enough to work from, and not something to take on faith. Transcription runs on Whisper Large V3 with Azure, summaries run on your choice of leading language models, and industry vocabulary cuts down errors on proper nouns and technical terms. The real safeguard is structural. Because a full transcript exists, you can check any line of any summary against what was said. A product that only gives you the summary offers you nothing to check it against.
Start with your next conversation
Apple entering this category is good news, and the split in the coverage is worth taking seriously on both sides. Sound Recognition is a genuine accessibility gain. Live Rewind will save you a small embarrassment now and then. Neither one is trying to leave you a record, and Apple has been straightforward about that.
So the question is not which product is better. It is what you need to have in your hands tomorrow morning. If the answer is a nudge, wait for the watch. If the answer is a transcript you can quote, action items with names attached, and a file you can still open next year, you need a device built for that job. The most wearable AI note taker in the Plaud lineup costs $179, includes 300 free transcription minutes a month, and starts working on your next conversation.
Live Rewind and Siri Recap are in beta and were announced on September 9, 2026. Feature details, availability, and regional support may change. If required by law, obtain consent from all participants before recording, and comply with applicable law.





















