I Fed a 45-Minute Meeting Recording Into Gemini Notebook — Here's What It Got Right (and Wrong)
I record most of my own work calls as voice memos — no fancy setup, just my phone on the table. They pile up unwatched the same way long YouTube videos do. So I ran the same test I did with a 3-hour YouTube course on an actual 45-minute meeting recording. Here's what turning raw audio into searchable notes and action items actually looks like.
Section 1: What Gemini Notebook Actually Does With an Audio File
Quick clarification before anything else, because two very similar-sounding Google features get mixed up constantly: Gemini in Google Meet can take live notes during a scheduled video call automatically. That's a real-time feature built into Meet itself. Gemini Notebook — the tool this post covers, formerly known as NotebookLM — does something different: you upload an existing audio file, from any source, and it becomes a searchable, question-answerable document inside a notebook. It doesn't need to have been recorded in Meet at all. A voice memo from your phone, a Zoom recording you exported, an in-person conversation you recorded on a handheld recorder — all of it works the same way, because Gemini Notebook only ever sees the audio file itself.
That distinction matters if your meetings happen across Zoom, Teams, in-person rooms, or a mix — Gemini Notebook doesn't care which platform the recording came from. It only cares about the file.
Here's what's technically supported when you upload audio as a source:
- A wide range of audio and video-with-audio formats are accepted, including MP3, WAV, M4A, AAC, OGG, and several others — you almost never need to convert a file before uploading it.
- The file needs actual speech in it. Ambient recordings, music without vocals, or long silent stretches aren't supported — there's nothing for the tool to transcribe.
- Each source can hold up to 200MB or roughly 500,000 words once transcribed, whichever limit hits first — enough for the overwhelming majority of meetings, interviews, or lecture recordings.
- Audio import supports dozens of languages, not just English, though transcription accuracy depends heavily on recording quality.
- Low-quality audio can cause the import to fail outright, or transcribe poorly. A phone recorded across a large conference room with cross-talk will give you noticeably worse results than a close, clear mic.
One honest practical note before you build a habit around this: if you're recording other people — a team meeting, a client call, an interview — check that you have their consent and that it's consistent with your company's policy and local recording laws before you upload it anywhere. That's true regardless of which AI tool you use afterward.
Section 2: My Real Test — a 45-Minute Meeting Recording, Start to Finish
I picked a recent internal planning call — 45 minutes, three participants, recorded as a single audio file on my phone, nothing special about the setup. This was a genuinely messy, real recording, not a clean studio test.
What I did
- Created a new notebook and titled it after the meeting's date and topic.
- Selected "Add," uploaded the audio file directly from my device.
- Waited for Gemini Notebook to process and transcribe the recording as a source.
- Opened the Source Guide for the auto-generated overview before asking anything myself.
- Asked chat directly for a structured list of decisions and action items.
Where it worked well
The auto-summary correctly separated discussion from decisions. Rather than just a chronological recap, the Source Guide picked out what had actually been decided versus what was still being debated — which, for a meandering planning call, was genuinely more useful than my own notes from the meeting.
Action items came out cleanly when I asked for them directly. A specific chat request for "list every action item mentioned, with who owns it" returned an accurate, correctly-attributed list. This is the single most useful practical output for a meeting recording — it turns a call nobody wants to re-listen to into something you can act on in under a minute.
It handled three overlapping speakers reasonably well. There were a couple of points where people talked over each other, and the summary still captured the gist of what was said, even if it occasionally attributed a comment to the wrong speaker.
Where it fell short
Speaker attribution wasn't fully reliable. On a call with people who don't have obviously distinct voices, expect the occasional mixed-up attribution — worth a quick read-through of anything you plan to act on or share, rather than trusting names blindly.
Background noise degraded quality noticeably. One stretch of the recording had a door open and hallway noise bleed in, and the transcript accuracy visibly dropped for that section. Cleaner audio produced meaningfully better output across the whole test.
Tone and context can get flattened. A slightly sarcastic aside in the meeting got summarized as a straightforward statement. If nuance or tone matters for a specific moment, that's a reason to go back to the audio itself rather than trust the text summary alone.
The honest verdict: for turning a meeting nobody wants to re-listen to into a usable record of decisions and action items, this is a genuine time-saver. For anything where exact wording, tone, or precise attribution matters — treat the summary as a strong first draft, not a transcript you'd forward without checking.
Section 3: Step-by-Step — How to Analyze Your Own Audio
Basic setup
- Open Gemini Notebook and create a new notebook, or add to an existing one if this recording relates to an ongoing project.
- Select "Add" to add a source, then upload your audio file directly from your device or Google Drive.
- Wait for processing. Gemini Notebook transcribes the file automatically — this can take a little longer for longer recordings or lower-quality audio.
- Open the Source Guide for the auto-generated overview before diving into specific questions.
- Ask targeted follow-up questions in chat — "what did we decide about the launch date," "list every action item and owner," "what concerns did anyone raise about the budget."
Getting better results
- Record close to the source. A phone placed centrally on a small table beats one across a large room every time — audio quality is the single biggest factor in transcript accuracy.
- Ask for structure, not just a summary. Requesting a specific format — decisions, action items with owners, open questions — produces something you can actually act on, rather than a paragraph you still have to parse yourself.
- Verify speaker attribution before sharing. Skim the output before forwarding it, especially on calls with several similar-sounding voices.
- Turn a recurring meeting into a running notebook. Add each week's recording as a new source in the same notebook, and you can ask questions across multiple meetings at once — "has this concern come up before" — instead of treating each recording as an island.
- Generate an Audio Overview of your own meeting notes. If you'd rather review a summary by listening than reading, the same source can produce a short two-host discussion of the meeting — the same feature covered in our guide to Gemini Audio Overviews for passive learning, applied to your own recordings instead of research material.
Who this is genuinely useful for
This is a strong fit if you regularly sit through calls that generate more talk than clear next steps — planning meetings, client calls, interviews, or recurring standups — and want a reliable written record without manually taking notes. It's a weaker fit for anything requiring word-for-word accuracy or legal-grade transcription, where you'd want a dedicated transcription service and a human review pass instead.
If you're stitching this into a broader workflow alongside video and document sources, our earlier post on converting YouTube videos into text summaries with Gemini Notebook covers the same tool applied to a different source type — the setup and limitations carry over almost exactly.
Do you record your own meetings, or rely on live transcription tools? Tell me in the comments which recordings you'd actually trust an AI summary of — and which ones you wouldn't.
No comments:
Post a Comment