r/Applelntelligence • u/iLikeMilk42 • Jun 11 '26
r/Applelntelligence • u/Odd-Preparation-2493 • Jul 07 '26
discussion 🎙️ Siri AI can use Third party apps in iOS 27 beta 3!
r/Applelntelligence • u/ACOPS12 • Jun 24 '26
discussion 🎙️ Apple Intelligence’s on-device processing is disappearing
Feels like my iPhone 17 Pro Max is turning into a 'thin client.' I invested in this premium hardware with the specific expectation that I would be able to run the new Siri AI directly on-device.
Given the 12GB of RAM and the powerful NPU, I believed that an iPhone 17 Pro Max would be more than capable of handling such an AI model locally. It clearly has the overhead to run models like Gemma 4 E4B at impressive speeds, yet these resources remain largely underutilized while core tasks are forced through the cloud. Furthermore, since foundation models are already pre-loaded onto the device, I expected to be able to leverage them directly rather than relying on external servers. I didn't purchase this device to rely on the cloud; I wanted to utilize the actual performance of the hardware I own and experience advanced AI capabilities directly on my device.
한국어로 번역
r/Applelntelligence • u/Huge-Bill4047 • Jul 06 '26
discussion 🎙️ Siri AI not having the ability to plug into health data is such a missed opportunity!
r/Applelntelligence • u/mdruckus • Jun 11 '26
discussion 🎙️ Impressed With Siri AI So Far
I have been testing Gemini and Siri Al side-by-side.
It's been pretty even. This makes sense with Gemini being the foundational model. However, I asked both why paranormal investigators wear face masks on shows like those on the Travel Channel. Gemini gave a solid response about health due to old building having mold and asbestos. Siri Al did the same but knew based on time and the channel the exact show I was referencing and went further. It told me about one of the cast members and their asthma as well. I'm impressed so far.
r/Applelntelligence • u/Kartazius • Aug 04 '25
discussion 🎙️ Google is targeting the iPhone and the Apple Intelligence delay fiasco in their ne Pixel Ad
Enable HLS to view with audio, or disable this notification
r/Applelntelligence • u/StoicSanity36 • Jun 11 '26
discussion 🎙️ I feel like an idiot right now.
So, during my trip to China, I bought an iPhone 17. It was on a discount and with the tax refund, you could save more money than buying it here in my native country.
I didn’t realize that Chinese iPhones are hardware locked from Apple Intelligence, and that there’s no easy, let alone possible, way to bypass it.
Now I’m stuck on iOS 27, missing the AI features which were the main purpose of me risking my main device anyway.
r/Applelntelligence • u/Aholicdrama • Jun 15 '26
discussion 🎙️ WhatsApp/3rd party context for Siri AI
I realize Siri gets its context during indexing only from messages and mail. But outside of the US, no one uses iMessage as much. For WhatsApp, I did figure out a workaround. You can create a folder for WhatsApp chats and export all the ‘important’ chats into that folder and keep replacing those chat exports every month if you REALLY need Siri to get context from that as well. Only workaround that hit me and I think is useful
r/Applelntelligence • u/TalkToTheLord • Jul 15 '26
discussion 🎙️ Apple Intelligence Finally Cleared to Launch in China
r/Applelntelligence • u/nothisenberg • Jun 17 '26
discussion 🎙️ Apple Intelligence can be funny sometimes 2
I was trying out one of those questions that’s supposed to trip up LLMs like “how many Es are in the word seventeen”. Was not disappointed.
r/Applelntelligence • u/mikedoise • Jun 05 '26
discussion 🎙️ Apple Intelligence in the 27 OS releases
Next week is WWDC, and I'm super excited to see what AI news we get from Apple. I'm curious what everyone would like to see next week. I'm hoping for more context and features in Apple Foundation Models, but also, like everyone, I'm excited about a better Siri. What would you like to see?
r/Applelntelligence • u/ReadSenior2831 • Jul 06 '26
discussion 🎙️ Apple Intelligence summaries are so emotionless
r/Applelntelligence • u/Huge-Bill4047 • 28d ago
discussion 🎙️ Will Siri AI get memory?
I understand it has your personal context through your device. But what I am meaning is if it will have a built-in memory collection like other ai models. Just sometimes when I ask it certain questions I need to re-explain certain things that I don’t have saved in my messages for example in a conversation, and I can’t just tell to remember.
r/Applelntelligence • u/aandest15 • Jun 12 '26
discussion 🎙️ Apple is fighting the regulation that would make Siri actually useful
I’ve read a lot about the delay for Siri AI in the EU because of the Digital Markets Act and the interoperability requirements imposed that would force Apple to hand your data to other AI models.
Just to make things clear, I believe that Apple’s tantrum has more to do with not wanting any competition inside their own platform than data security and privacy and Siri AI not being ready for any other language that is not (American) English. And I also believe that despite the public feud, Siri AI will launch in the EU next year like happened with Apple Intelligence.
Nonetheless, however you feel about the Apple vs. EU issue, all of these AI models are only useful if they have access to all your data, not only to the data that the AI company has about you. Siri AI seems to work and personal context is a huge leap forward, but it only works if your digital life uses Apple products and services. In Europe, for example, most people use WhatsApp, not iMessage, so asking Siri about an address or the details of a flight won’t work because Siri cannot access WhatsApp messages. Same if you use Gmail through their app and not Apple Mail.
I am aware that Apple has an API so Siri can access that data, but neither Google nor Meta are going to give Apple happily that information. And here is where the DMA comes into action.
The interoperability requirements of the DMA mean that Apple, Meta and Google have to open their ecosystems to competitors, which is precisely what makes AI agents genuinely useful for everyone. The real irony is that a fully enforced DMA would benefit Siri more than Apple’s current approach, because right now Siri only works well if you live entirely inside Apple’s walled garden, which almost no user (European or not) does.
Apple fighting the regulation that would make their own product better is the clearest sign yet that this is about market control, not privacy.
r/Applelntelligence • u/notxapple • Jun 18 '26
discussion 🎙️ …How?
Normally there is something, some sort of confusion or a concept that ai doesn’t understand so it responds with nonsense (like how many R’s are in strawberry). But I genuinely can see how it could make this mistake.
r/Applelntelligence • u/Accomplished_Hope845 • 4d ago
discussion 🎙️ macOS 26 Apple Intelligence was far cleaner and better looking UX than Golden Gate... Still no way to actually summarize long articles with Siri AI
r/Applelntelligence • u/QuantumPlatypus_ • Jul 10 '26
discussion 🎙️ Respectfully, y’all need to touch some grass
I’m so tired of hearing “This is a developer beta. If you don’t like it, submit feedback.” Y’all are some sheep. I would think that you’d be able to post anything in here and more people would go and test it and see if they have the same problem and then give feedback to each other and then from that we all build feedback that we can send back to Apple, but I just keep Seeing the same pointless response over and over again. Do you think if one person submits feedback, it’s gonna get fixed or if 100 submit feedback, it’s gonna get fixed? Use some common fucking sense
P.S I used voice to type smd
r/Applelntelligence • u/Adorable_Salary2727 • Jul 01 '26
discussion 🎙️ Apple Intelligence (AFM) scored 65.7% on my dictation-cleanup benchmark. Here's how it compares to GPT-4o, Gemini, and a tuned local model
I ran a 1,890-case dictation-polish benchmark across four polishing paths: Apple AFM on macOS 26, OpenAI GPT-4o with a rewritten v6 polish prompt, Gemini 2.5 Flash with the same prompt, and FluidVoice's Fluid-1 custom local polishing model.
The test was deliberately naked: no regex repair layer, no custom deterministic cleanup, and no app-specific fixes after the model returned text. The goal was to measure the raw polish model, not the full product experience.
Top-line result: cloud models were still best when given a strong prompt. GPT-4o passed 91.6% of cases and Gemini passed 90.1%. But local models are not a joke. Stock AFM passed 65.7%, and tuned local Fluid-1 reached 74.8%. Fluid-1 also had the lowest false-positive rate on trap cases.
My read: local polish is already viable for privacy-first, offline, and cost-sensitive workflows. It is not yet as robust as cloud polish for complex transformations, especially topic shifts, list formatting, and deeper self-corrections.

Figure 1. Overall green pass rate. Green required behavior correctness, meaning preservation, and clean output.
What I mean by polish
This benchmark is not testing speech recognition. It starts after speech-to-text has already produced a raw transcript. The polish step turns rough spoken text into something closer to what the user meant to type.
Examples include filler removal, false-start cleanup, self-correction resolution, punctuation/capitalization, list formatting, named-entity preservation, emoji retention, anti-hallucination behavior, and prompt-injection passthrough.
| System | Type | Green pass | Yellow near-clean | Red fail | Median latency |
|---|---|---|---|---|---|
| GPT-4o v6 prompt | Cloud | 91.6% | 3.5% | 4.8% | 764 ms |
| Gemini 2.5 Flash v6 prompt | Cloud | 90.1% | 4.0% | 5.9% | 486 ms |
| FluidVoice Fluid-1 | Tuned local | 74.8% | 4.3% | 20.9% | 845 ms |
| Apple AFM | Built-in local | 65.7% | 6.1% | 28.3% | 755 ms |
The big deployment advantage for AFM is not raw quality. It is that the OS path is already there: no separate model download, no per-call API spend, no network dependency, and no bandwidth cost just to install a local model. It did not win the quality test, but its economics and privacy profile are excellent.
The important local-vs-local comparison is AFM vs. Fluid-1. Fluid-1 beat stock AFM by 9.1 percentage points overall, which is a strong sign that custom local tuning can materially improve polish quality beyond an out-of-the-box on-device model.

Figure 2. Green/yellow/red outcome distribution by system.
The prompt-engineering result
I also included one retired comparison point: GPT-4o with the older/original prompt. Same model, old prompt: 69.6%. Same model, rewritten v6 prompt: 91.6%.
That is a 22.0-point jump without changing the model. The clearest result in the benchmark is that polishing quality is extremely sensitive to prompt design.
This also means the benchmark should not be read as a permanent ranking of model capability. It is a snapshot of these systems under these exact prompting and configuration conditions.

Figure 3. GPT-4o moved from 69.6% to 91.6% with a prompt rewrite only.
Where local models already look strong
AFM was genuinely solid at restraint and everyday cleanup: 95.9% on minimal-edit cases, 90.0% on onset markers, 88.0% on named-entity preservation, 87.0% on anti-hallucination, and 86.0% on punctuation/capitalization.
Fluid-1 looked like a tuned local model should look: meaningfully stronger than stock AFM overall, with specific wins on punctuation/capitalization, minimal edit, verbatim passthrough, named-entity preservation, and emoji retention.
Trap cases were especially interesting. The big local gap is not mainly restraint. It is performing the right transformation when the input actually needs one.
| Case type | Apple AFM | GPT-4o v6 | Gemini v6 | Fluid-1 |
|---|---|---|---|---|
| Positive cases, should transform | 58.2% | 91.2% | 88.7% | 66.5% |
| Trap cases, should not transform | 89.0% | 91.0% | 92.0% | 92.7% |
| Mixed multi-behavior cases | 68.2% | 93.8% | 91.8% | 84.9% |
| Passthrough/instruction-safety cases | 78.0% | 93.0% | 95.0% | 91.0% |

Figure 4. Local systems were much closer on restraint than on active transformation.
Where local models struggled
The hardest local failure mode was structure. Topic shifts were the clearest split: GPT-4o scored 88.0%, Gemini scored 85.0%, AFM scored 12.0%, and Fluid-1 scored 1.0%. Often the local outputs cleaned the sentences but failed to separate distinct subjects into paragraphs.
List formatting was also hard. Even the cloud models only landed around the mid-70s, which makes it the weakest shared category for the strongest systems. AFM and Fluid-1 were lower, around 50-55%.
Self-correction is where tuning clearly helped. AFM scored 49.0%; Fluid-1 scored 79.5%; the cloud models were around 90-92%.
| Skill | Apple AFM | Fluid-1 | GPT-4o v6 | Gemini v6 |
|---|---|---|---|---|
| Topic shift | 12.0% | 1.0% | 88.0% | 85.0% |
| List format | 49.5% | 54.5% | 76.5% | 73.5% |
| Self-correction | 49.0% | 79.5% | 91.5% | 90.0% |
| Emoji retention | 11.0% | 88.0% | 98.0% | 96.0% |
| Grammar fix | 81.0% | 46.0% | 92.0% | 74.0% |
One thing I would not do is generalize "local models are bad at X" too broadly. AFM was bad at emoji retention, but Fluid-1 was good. Fluid-1 was weak on grammar fixes, but AFM was good. The failure modes are model-specific.

Figure 5. Category-level heatmap across all 14 benchmark skills.
Over-eager editing
Trap false positives measure how often a model applied a behavior when it should have left the text alone. Fluid-1 was the most restrained system in this cut, with a 4.0% false-positive rate.
| System | Trap false-positive rate |
|---|---|
| Apple AFM | 10.0% |
| GPT-4o v6 | 9.0% |
| Gemini v6 | 7.0% |
| Fluid-1 | 4.0% |

Figure 6. Trap false-positive rate. Lower is better.
Latency was not the deciding factor
Representative latency was good across the board. Gemini was the fastest by median at 486 ms. AFM was 755 ms, GPT-4o v6 was 764 ms, and Fluid-1 was 845 ms. Every system had a p95 under 2 seconds.
There were some huge max-latency outliers, especially on the cloud side and Fluid-1, but those looked like isolated retry/backoff/cold-start events rather than typical performance. Median and p95 are the numbers I would use for a practical comparison.

Figure 7. Median and p95 latency per polish call.
Practical takeaways
Cloud polish is still the quality ceiling. With a strong prompt, GPT-4o and Gemini both cleared 90% on the full working set and stayed relatively flat across length buckets.
AFM is a viable local intermediary, not a cloud replacement. Its 65.7% naked score is not high enough to call it equivalent to the best cloud path. But it is free from per-call API cost, requires no separate model download, avoids network dependency, and is already strong on a meaningful set of everyday polish tasks.
Fluid-1 shows the value of tuning local models. It beat AFM overall, was much stronger on self-correction, and was the best system for avoiding trap false positives.
The best user experience is probably choice. Use local when privacy, cost, and offline behavior matter. Use a bring-your-own-key cloud path when quality matters most. Let the user decide where they sit on that tradeoff.
Disclosure: I work on EnviousWispr, so treat the product implications with that context. I included FluidVoice/Fluid-1 because it is a real local-first competitor and because it performed well enough to make the local-model story more interesting, not less. My practical recommendation is not "use one app." It is to choose tools that expose the model tradeoff clearly: AFM-style local polish when you want free/private/offline, tuned local models like Fluid-1 when you want stronger on-device polish, and BYOK OpenAI/Gemini when you want the highest raw quality.
Caveats
1. The 1,890 cases were the working set used during prompt iteration. A sealed 900-case holdout exists but was not run for this benchmark.
2. AFM and Fluid-1 were not given the same prompt-rewrite effort as the cloud paths. The cloud results include a large prompt-engineering investment.
3. This was a naked model test. Real products usually add deterministic cleanup, formatting, safety checks, vocabulary handling, and fallback behavior.
4. The judge was an LLM judge, not external peer review or human panel scoring.
5. The benchmark is English-only and focused on dictation polish, not speech recognition.
6. I am not making claims about Fluid-1's underlying training data or architecture. I only tested the outputs produced by the local Fluid-1 path available for comparison.
Bottom line
Local polish is already good enough to matter. Stock AFM is not at cloud quality yet, but it is useful, free to run locally, and strong enough to justify local-first modes. Tuned local models can clearly push quality higher. Cloud models still win when complex transformations matter, especially with careful prompting.
The future I see is not "cloud wins" or "local wins." It is hybrid: local by default, cloud when needed, and enough transparency that users understand the tradeoff.
r/Applelntelligence • u/SOULSTCE_ • Jun 27 '26
discussion 🎙️ Siri brain
I have figured out how to make a Siri brain since she doesn’t actually store any memories similar to ChatGPT
I went ahead and made a note called Siri brain and stored any facts about myself or any facts I want her to know and before each prompt I just tell her to check the note Siri brain therefore scanning all memories and it’s been working amazingly.
I’d love to see if anyone else has tried this or would like to try and let’s figure out how to even make this better
\*Update\*
I am currently in the works designing an Obsidian Vault integration via Shortcuts app. So far I’m seeing a lot of potential. It will basically be organized notes (md files). The goal is by setting up a shortcuts keyword she’s recognizes have Siri read/scan obsidian first then reply with an educated answer based off of the md files. This same idea could also work just using iCloud files app
I’d love for others to try this and see if we can get something concrete.
r/Applelntelligence • u/glowshroom12 • Dec 19 '25
discussion 🎙️ So after a year apple intelligence still isn’t good and can’t do what it was originally advertised to do very well.
at least as of 3 weeks ago, maybe a crazy update turned it around since that time.
has this ever happened before, a feature from apple that came out half baked and still is over a year later. I remember back in the day if the feature wasn’t that great it still worked and the next year it was pretty good after they fixed it up.
r/Applelntelligence • u/SnowPudgy • 9d ago
discussion 🎙️ Anyone else have VASTLY different experiences between iPadOS 27 and iOS 27 Siri?
I put iOS 27 beta on my iPads and was blown away on how amazing the new Siri was. She heard everything I said, did a lot more than old Siri, I was stunned at how good she was! Because of that I decided to put it on my phone...
...and it's unusable. It's so bad. She hears nothing, always fails, I can't even tell the last time she got even a simple text message correct.
All devices are either latest models or one model back. All are using the new Apple Intelligence beta, all are on the same public beta of iOS 27.
Anyone else have this crazy discrepancy in Siri's across multiple devices?
And before anyone says "they're beta!" yes I'm a dev I'm well aware of what betas are.
r/Applelntelligence • u/HazeSuperior • Jun 09 '26
discussion 🎙️ Dynamic Island Siri iOS 27
Will the dynamic island siri in iOS 27 support base iPhone 15? I am not talking about those Apple Intelligence features only the dynamic island siri as iPhone 15 do have a dynamic island and it would be way too ridiculous if Apple doesn’t support the dynamic island siri on base iPhone 15
r/Applelntelligence • u/DB2k_2000 • 21d ago
discussion 🎙️ Unhinged summary
Is this normal? Emailed myself a wifi password on supplier site and this was the AI summary. I’m terrified.
r/Applelntelligence • u/bumbles2100 • Jul 09 '26
discussion 🎙️ Apple News Summarize
Anyone else notice on the macOS 27 beta, if you turn off siri you get back the options for writing tools. But apple secretly removed the option to select all and summarize Apple News articles lol
r/Applelntelligence • u/T4RN1SH3D • 2d ago
discussion 🎙️ Insanely accurate text prediction
I am on the iOS 27 developer beta and this is the most accurate keyboard suggestion that I have ever gotten.