Verdict: Google's mid-2026 Gemini update is the most aggressive agentic push from any major AI lab this year — Lyria 3.5 turns text into full songs, Gemini Spark now drives your Chrome browser with your saved logins, Speak to Window puts voice-controlled AI on top of every macOS app, and scheduled actions let Gemini run tasks while you sleep. For anyone already paying for Google AI Pro ($19.99/month), most of these features are included. The catch: several are macOS-only or US-only at launch, and the browser automation introduces real security considerations that Chrome Enterprise has already shipped a kill-switch for. Here's what each feature does, how to set it up, and whether it's worth your time.
Last verified: 2026-08-05 · 5 major features covered · Pricing volatile — recheck monthly
- Best for automation: Gemini Spark + Chrome auto-browse (US, AI Pro required)
- Best for content creators: Lyria 3.5 music generation (free tier available)
- Best for hands-free work: Speak to Window voice dictation (macOS only)
- Best for hands-off ops: Scheduled actions (AI Pro/Ultra, 10-task cap)
- Pricing: Google AI Pro $19.99/mo · Google AI Ultra $99.99/mo
What's actually new in Google Gemini's 2026 update?
Google's 2026 Gemini update isn't one release — it's a wave of features rolled out between April and August 2026 that collectively transform Gemini from a chatbot into an ambient AI layer that lives across your entire digital workspace. The five features that matter most are: Lyria 3.5 AI music generation, Gemini Spark Chrome integration, Gemini Skills, scheduled actions, and Speak to Window voice control.
Each addresses a different bottleneck: creating original audio content, automating browser tasks, codifying repeatable workflows, running tasks on a schedule, and controlling AI without switching apps. We've tested the setup path for each and verified the claims against Google's own announcements and documentation.
Can Gemini really make music now with Lyria 3.5?
Yes — Google DeepMind's Lyria 3.5 model generates full songs up to 3 minutes long from a text prompt, and you don't need a paid subscription to try it. The model card, published July 29, 2026, describes Lyria 3.5 as a latent-diffusion music generation system that takes text as input and produces audio plus lyrics as output.
You access it through Flow Music (at flowmusic.google), where you sign in with a Google account. From there, you can:
- Type a text prompt — describe the genre, mood, instruments, and vibe (e.g., "an upbeat electronic intro for a tech podcast, energetic and modern")
- Upload an image to inspire music — Lyria converts the visual mood into a track
- Specify vocals, lyrics, and structure — control verses, choruses, bridges, and vocal style
- Generate multiple options — pick your favourite from several generated tracks
- Add lyrics — ask it to write words matching the existing beat, then layer vocals on top
Lyria 3.5 also supports a music video mode — you provide a song, a main character image, and a style, and it generates a video up to 2 minutes long.
For developers, the underlying Lyria 3 API is available through Google's Gemini API with two models: lyria-3-clip-preview (30-second clips, best for loops and previews) and lyria-3-pro-preview (full-length songs with verses, choruses, and bridges). Output is 44.1 kHz stereo audio in MP3 or WAV format.
Limitation: Lyria 3.5's API documentation notes that multi-turn editing — refining a generated clip through follow-up prompts — is not supported in the current version. Each generation is single-turn.
Sources: Google DeepMind Lyria 3.5 Model Card (July 2026) · Google Blog: Introducing Lyria 3.5 · Google AI for Developers: Lyria 3 API docs
What does Gemini Spark's Chrome integration actually do?
Gemini Spark's Chrome integration, announced July 30, 2026, lets Google's AI agent operate inside your actual Chrome browser — using your logged-in accounts and saved passwords to complete multi-step web errands on your behalf. This is a significant shift from Spark's earlier remote-browser approach, which ran in an isolated cloud instance without access to your personal sessions.
For a deeper look at whether Gemini Spark on Google AI Pro is worth the subscription, see our Gemini Spark decision guide after the August global rollout.
With your permission, Spark in Chrome can now:
- Book travel by comparing flight options and starting the checkout flow
- Schedule apartment viewings from listings you've saved
- Make restaurant reservations on sites where you're signed in
- Log into loyalty programs using credentials from Google Password Manager
- Research with cited sources and create custom news digests
How to set up Chrome auto-browse
- Open Gemini Spark in the Gemini app (web or desktop)
- Navigate to Settings > Connected Apps and enable Chrome browsing
- Grant permission for Spark to use your Chrome session
- Describe your task in natural language (e.g., "Find flights from JFK to SFO under $400 for next Friday and start the booking")
- Spark shows you each step; you verify before sensitive actions like payments
Safety guardrails
Google built in several protections: payment handoffs require manual confirmation, sensitive operations need user verification, and the system includes prompt-injection screening. Chrome Enterprise also shipped a dedicated policy to block the feature organisation-wide — a signal that Google itself considers the capability significant enough to warrant an IT kill-switch from day one.
Availability
| Feature | Where | Plan Required | Platform |
|---|---|---|---|
| Gemini Spark (core) | US + 160+ countries | Google AI Pro or Ultra | Web, Mac, mobile |
| Chrome auto-browse | US only (for now) | Google AI Pro or Ultra | Desktop Chrome only |
| Remote browser tasks | Wherever Spark is available | Google AI Pro or Ultra | All Spark platforms |
Sources: Google Blog: Gemini Spark Chrome integration (July 30, 2026) · Google Support: What's new for Gemini Spark
How do MCP connected apps work with Gemini Spark?
MCP (Model Context Protocol) is an open standard that lets AI models connect to external tools through a standardised bridge. Gemini Spark supports MCP connections, which means you can link it to services beyond Google's own ecosystem.
To connect an MCP tool:
- Open Gemini Spark > Connected Apps
- Scroll to the MCP section
- Add the tool you want to connect (it must support the MCP standard)
- Once connected, Spark can read from and act on that tool's data
Real-world examples: connect your email tool and calendar, then ask Spark to "scan my inbox for anyone asking about our service and draft a warm reply explaining how it works." Or connect a community platform and ask Spark to "review our landing page and suggest three improvements to the wording."
Google has been rolling out third-party Connected Apps throughout July 2026 — Canva, Instacart, OpenTable, Dropbox, and Zillow support were announced in July alone. For the broader picture of how Gemini's agentic features evolved through 2026, see our earlier breakdown of Gemini's 2026 agentic workbench updates.
Source: Google Support: What's new for Gemini Spark (July 2026)
What are Gemini Skills and how do you create one?
Gemini Skills are reusable instruction sets that tell Gemini how to handle a specific type of task every time, so you don't have to re-explain your preferences. Think of them as custom Standard Operating Procedures that Gemini follows automatically.
A skill can be:
- Always write in your brand's voice — set the tone, vocabulary, and style once
- Generate three content ideas every week — a recurring instruction Spark follows on schedule
- Format every email response with your signature and disclaimer — applied automatically
- Summarize any article in exactly five bullet points — a consistent output format
Skills persist across conversations, so Gemini remembers them without you repeating the instructions. This is distinct from scheduled actions (which run on a timer) — skills define how Gemini does something, while scheduled actions define when.
Source: Google Support: Use Gemini Spark to manage tasks & workflows
Can you schedule Gemini tasks to run automatically?
Yes — Gemini Scheduled Actions let you set tasks that run on a timer, react to events, or watch for changes and notify you. The feature is available to paid subscribers (Google AI Pro or AI Ultra) and qualifying Google Workspace plans.
What you can schedule
- Morning briefings — "Every weekday at 8 AM, summarize my unread emails and today's calendar"
- Monitoring — "Watch for new reviews about my business and notify me when one appears"
- Recurring research — "Every Monday, find three news stories about [topic] and summarize them"
- One-off scheduled tasks — "At 5 PM today, draft a weekly report summary"
How to set up a scheduled action
- Open the Gemini app (mobile or web)
- Type your prompt with a timing instruction (e.g., "Every morning at 7 AM, check my calendar and give me a summary of today's meetings")
- Gemini confirms the schedule and asks you to review it
- Connect any required apps (Gmail, Calendar) if prompted
- Manage your actions in the Scheduled Actions manager tab
Limitations
- Maximum 10 active scheduled actions per account
- Inactivity cutoff — Gemini may pause actions you haven't engaged with in a while
- Available on mobile and web only — not deeply integrated into Gmail or Docs
- Requires a paid plan — free tier users don't get access
Compared to ChatGPT's equivalent task feature (also capped at 10 tasks, same $20/month price point), Gemini's advantage is native Google Workspace integration — it can pull your actual emails, calendar, and docs rather than producing generic outputs.
Sources: Google Support: Gemini Spark task management · Digital Citizen: Gemini Scheduled Actions guide (April 2026)
What is Speak to Window and how does it work?
Speak to Window is a voice control feature for the Gemini desktop app (macOS) that lets you issue commands to Gemini from inside any application, without switching tabs. You hold the Fn key, speak, and Gemini transcribes or acts based on your voice — with context from whatever's on your screen.
Announced July 29, 2026, and rolling out globally in English on macOS, Speak to Window works like this:
- Hold the Fn key (the same key Apple uses for system dictation)
- Speak your command — "Write an email to my team about tomorrow's meeting agenda"
- Gemini processes your intent — it doesn't just transcribe; it understands what you want and acts
- Output appears at your cursor — in the active app, not in a separate Gemini window
What makes it different from regular dictation
Standard macOS dictation converts speech to text. Speak to Window processes the intent behind your words and can write, summarize, edit, or generate based on what's on your screen. With the optional reasoning toggle enabled, Gemini thinks deeper before responding and can reference whatever's visible on your display.
Practical examples
- While writing an email: "Make this more persuasive and add a call to action"
- While reading a document: "Summarize this in three bullet points"
- While working in a spreadsheet: "Format this data as a comparison table"
- While browsing: "Highlight this article and give me the key takeaways for my business"
Limitations
- macOS only at launch (Windows not yet supported)
- English only initially, with more languages planned
- The Fn key activation is a deliberate choice — it mirrors Apple's own dictation shortcut, positioning Gemini as a smarter replacement
Magic Pointer, a companion feature, combines screen pointing with voice prompts — you point at something on your display and ask Gemini about it, getting a response that understands what you're looking at.
Sources: Google Blog: Gemini for macOS natural language capabilities · Gemini release updates (July 29, 2026) · TestingCatalog: Speak to Window rollout
How to use Gemini inside any Chrome web page
Google also rolled out in-page Gemini integration for Chrome. When you're reading any web page, you can highlight text and ask Gemini to summarize it, explain it, or give you examples specific to your own business. It reads the full page context and delivers an answer built around what you're doing — without leaving the page.
This is part of the broader Gemini-in-Chrome push that includes the Spark auto-browse feature above, but in-page Gemini is lighter weight: it's about reading and analyzing content rather than taking actions on your behalf.
How much does Google Gemini cost in 2026?
| Plan | Price | Key Features | Who It's For |
|---|---|---|---|
| Free | $0 | Gemini Flash, limited usage, voice mode, some music generation | Casual users |
| Google AI Pro | $19.99/mo | Gemini 3 Pro access, Spark, scheduled actions, Chrome auto-browse (US), Workspace integration, 5TB storage | Professionals, small business |
| Google AI Ultra | $99.99/mo | Highest usage limits, priority access, all Pro features, expanded credits | Power users hitting Pro caps |
Pricing restructured in May 2026 — the old "Gemini Advanced" branding was replaced with the current Google AI tier system. The free tier now includes Flash model access and limited feature trials. Most agentic features (Spark, scheduled actions, Chrome auto-browse) require AI Pro or Ultra.
Volatile fact: Google has changed Gemini pricing multiple times in 2026. Verify at one.google.com before subscribing.
Sources: Costbench: Google Gemini Pricing (verified May 2026) · Google One pricing page
What this means for you
If you run a small business: The highest-ROI combination is Gemini Spark + Chrome auto-browse (for automating repetitive web tasks like booking, research, and data entry) + scheduled actions (for daily briefings and monitoring). This stack replaces several manual hours per week for under $20/month — but only if you're in the US and on macOS/Chrome desktop. For a broader framework on wiring AI agents into your operations, see our guide on building an agent OS for your business.
If you're a content creator: Lyria 3.5 is the standout. Full songs with vocals, lyrics, and music videos — at no cost through Flow Music — eliminate the need for royalty-free music libraries or hired composers for basic needs. The 3-minute cap and single-turn editing limitation mean it won't replace professional production, but for intros, background tracks, and social content, it's more than enough.
If you work hands-free: Speak to Window is the most genuinely novel feature for anyone who thinks faster than they type. Being able to say "summarize this" while reading a long document, or "draft a response" while looking at an email — without leaving your current app — is a workflow upgrade that competitors haven't matched at this depth.
If you're evaluating vs ChatGPT: Gemini's agentic features (Spark, Chrome integration, scheduled actions) go further than ChatGPT's current offerings for browser-based automation. But ChatGPT still leads on desktop-level computer use and raw coding performance. The right answer for most teams is a multi-model approach — see our guide on building a multi-model AI coding workstation for the full strategy.
FAQ
Q: Is Google Gemini's Lyria 3.5 music generation free to use? A: Lyria 3.5 is accessible through Flow Music (flowmusic.google) with a Google account. The free tier of Gemini includes some music generation access. The underlying Lyria 3 API is pay-per-use for developers. Exact free-tier limits may change — verify at Google's Lyria page.
Q: Does Gemini Spark's Chrome auto-browse work outside the US? A: As of the July 30, 2026 announcement, Chrome auto-browse is rolling out first in the United States. Google said expansion to additional regions is planned but has not given a timeline. The remote browser (non-Chrome) version of Spark works in all 160+ countries where Spark is available.
Q: Can Speak to Window control any application on my Mac? A: Speak to Window works across any desktop application by holding the Fn key to issue voice commands. It processes intent (not just transcription) and can write, edit, summarize, or generate text at your cursor. It requires the Gemini desktop app for macOS and is rolling out globally in English.
Q: How many scheduled actions can I have active at once in Gemini? A: Gemini caps scheduled actions at 10 active tasks per account. This matches ChatGPT's equivalent limit. Scheduled actions are available to Google AI Pro and AI Ultra subscribers, and qualifying Google Workspace plans — not the free tier.
Q: What's the difference between Gemini Skills and Gemini Scheduled Actions? A: Skills define how Gemini should do something (your brand voice, output format, tone). Scheduled Actions define when Gemini should do something (every morning, when a review appears, on a recurring weekly cycle). Skills persist across all conversations; scheduled actions run independently on a timer.
Q: Is Google AI Pro worth $19.99/month for these features? A: For most professionals and small business owners already using Google Workspace, yes. The combination of Spark automation, scheduled actions, Workspace integration, and 2TB+ of storage at $19.99/month is competitive with ChatGPT Plus ($20/month). Upgrade to AI Ultra ($99.99/month) only if you consistently hit Pro's usage limits.

Discussion
0 comments