The voice-of-customer research you don’t have to earn
Get voice-of-customer insights without the lengthy research. Discover what your customers want and make smarter decisions.
Every piece of VoC advice you have ever read assumes something it never says out loud: that you already have the calls.
Mine your discovery calls. Tag your demos. Run your support tickets through a model and surface the recurring themes. All good advice, and all of it useless to a team that does not yet have the volume, or that has the volume locked in a CRM nobody has integrated, or that is trying to enter a category where they have run eleven discovery calls total.
So most teams skip it and go back to writing from internal assumptions.
Meanwhile there is a corpus of your exact buyers, on camera, unprompted, describing the problem in their own words. It has been accumulating for years. It is public. And almost nobody in demand gen reads it, because it is video and video cannot be read.
What is actually sitting there
Start with the obvious one and keep going.
Competitor user conferences. Every vendor of any size runs one and posts the sessions. The keynote is marketing. The breakout sessions are not. A breakout session is a customer standing on stage explaining, in operational detail, what was broken before they bought and what they measure now.
Customer webinars. The joint vendor-customer webinar is a sales asset, and it is also forty minutes of a practitioner answering unscripted audience questions at the end.
Case study videos. Short, heavily produced, low signal in the narration. The signal is in the interview segments where somebody describes their old process.
Analyst and industry panels. Four practitioners disagreeing about an approach in front of a room. This is where you find the objection language, because that is the room where objections get said out loud.
Podcast appearances by your buyers. Not vendor podcasts. The ones where your ICP goes on a peer show and talks about their job for an hour without a single person from marketing in the room.
That last category is the richest and the least mined, and it is entirely made of people who do not know they are doing research for you.
Getting it into text
The reason this corpus goes unused is not that people don’t know it exists. It’s that watching forty webinars is not a plan.
Reading forty transcripts is.
Vomo is the tool referenced in the steps below. Any of the credible options will do the same job; what matters for a research sprint is that you can process a two-hour session without hitting a length cap partway through.
Copy the video URL. Address bar. Unlisted works too, which covers a lot of gated webinar replays that ended up on an unlisted link.
Paste it and click Generate Transcript, It’s Free. No account required at this point.
Confirm the video it pulled. Title and duration come back. Ten seconds, and it saves you processing the wrong session off a conference playlist.
Click Transcribe. There is no cap on the length of a single video, which matters here more than anywhere, because conference sessions run long and most free tools quietly truncate them.
Read the summary first. You get timestamped chapters alongside the raw text.
Sign in to export. Free account, then TXT, DOCX, PDF or SRT.
Ten sessions is under an hour of processing and about ninety minutes of reading. That is a research sprint, not a project.
Four things to extract, and nothing else
Do not read these looking for insight in general. You will find nothing and conclude the method doesn’t work. Read for four specific artefacts.
Objection language. Not the objection. The wording. When a practitioner on a panel says “we looked at that and it felt like a lot of change management for something that touches four people,” that is a sentence you can answer directly in a landing page. Your internal version of that objection is “concerns about adoption,” which is answerable by nobody.
The real comparison set. Ask any prospect what they evaluated and you get the polite answer. Listen to a customer talk to peers and you get the actual one, which frequently includes “a spreadsheet,” “the thing we built internally,” and “nothing, we just kept doing it manually.” Those three are usually your largest competitors and none of them show up in a win-loss report.
Trigger events. The thing that happened right before the problem became urgent. New head of department. Failed audit. A big customer asked for something they couldn’t deliver. These are what your campaign timing should be built on and they are almost never in your persona doc.
Vocabulary, verbatim. This is the one that changes your copy the same week.
The vocabulary gap is wider than you think
Run a search across your ten transcripts for the term your category uses to describe itself.
It will barely appear.
Buyers do not say “workflow orchestration.” They say things are falling through the cracks. They do not say “single source of truth.” They say we have four spreadsheets and nobody knows which one is current. They do not say “accelerating time to value.” They ask how long until it actually does something.
You already knew this in the abstract. Reading forty transcripts makes it specific, and specific is the part you can use. You end up with a list of maybe fifteen phrases your market genuinely uses, ranked by how often they came up, and that list rewrites your headlines, your ad copy, your SEO targets and your sales one-pager without you needing to run a single interview.
A YouTube video transcript tool is not doing anything clever here. It is making a searchable corpus out of material that was already public. The insight was always there. It was just stored in the one format that can’t be searched.
Entering a market you can’t hear
The version of this that pays for itself fastest is market entry, and it is the one nobody attempts.
You are taking a product into Germany, or Japan, or Brazil. Your research consists of a localisation vendor, a competitor’s translated homepage, and whatever your one regional hire remembers. Meanwhile that market runs its own conferences, posts its own sessions, and has its own practitioner podcasts, all of which are describing the exact same problem in the local vocabulary that your translated copy is about to get wrong.
Transcription handles roughly 50 languages, so that corpus stops being closed. Convert the local sessions, read them, and find out what the category is actually called there, which is regularly not the translation of what it is called here.
Two cautions. Machine transcription of a second language plus machine translation compounds error, so anything you intend to publish needs a native speaker to check it. And the local comparison set will contain regional vendors you have never heard of, which is uncomfortable and is precisely the finding you needed before committing a budget.
Where this will mislead you
Four failure modes, and the first one is severe enough to disqualify the method if you ignore it.
Selection bias, badly. Nobody puts an unhappy customer on stage at their own user conference. Everything in this corpus is filtered through a vendor’s willingness to publish it, which means you are reading the best case, described by the most successful adopter, coached by a marketing team. Use it for language. Do not use it to size satisfaction.
Staged questions. The audience Q&A at the end of a joint webinar has usually been seeded. You can often tell which ones, because the answer arrives too fast and too complete. Weight the awkward questions higher.
Age. A conference session from 2022 describes a market that has since changed. Check the upload date before you build a campaign on it.
Transcription error on exactly the words you care about. Accuracy runs around 95% on clean audio across roughly 50 languages, and the remaining slice is concentrated in proper nouns: product names, vendor names, acronyms, job titles. Your competitor’s product name will come back misspelled at least once per session. Before you quote anything externally, verify it against the video at the timestamp.
Multi-speaker panels also need a pass by hand. Speakers get labelled automatically and you can rename them afterwards, but four people on a panel interrupting each other will produce attribution you have to fix before the transcript is quotable.
Run it as a two-week sprint
Week one is collection and conversion. Build a list of twenty sessions across the five categories above, weighted toward peer podcasts and breakout sessions rather than keynotes. Convert all twenty. Read the summaries and cut the list to the ten that are actually practitioners talking rather than marketing talking.
Week two is extraction. Read the ten properly and pull every instance of the four artefacts into one document, organised by artefact rather than by source. Objection language in one place. Comparison sets in another.
What you have at the end is a messaging brief built from your market’s own sentences rather than your team’s. It cost you two weeks of one person’s partial attention and no research budget.
On tooling cost, so it doesn’t derail the plan: the free tier covers 30 minutes of transcription a week, which is roughly one conference session. That is fine for testing the method and useless for the sprint. Unlimited is $1.92 a week, and the entire two weeks costs less than a single sponsored LinkedIn post.
The point
The assumption you have been operating under is that customer research requires access to customers.
It requires access to customers talking. Those are not the same constraint, and one of them has been solved in public for a decade while B2B marketing teams wrote personas from memory.


