Blog
July 31, 2026
 
·
 
Georgina Ford

Audio Visual & Text Filtering: Why the Source of a Mention Changes Everything

Text-First, API-Limited Tools Are Built for a World That No Longer Exists

Imagine this: three mentions of your brand pop up at the same time. One is a hashtag tucked into a caption. Another is someone mentioning you out loud, deep into a podcast episode. The third? Your logo flashes for just a couple of seconds in the background of a TikTok. Most monitoring tools will spot one and miss the other two completely. And even if they do catch all three, they usually treat them like they're exactly the same.

But not all mentions are created equal. Where a mention comes from, and how it’s shared, can totally change what it means for your brand.

Why Format Isn’t the Same as Signal

Most social listening tools were designed for a world where everything happened in text: captions, hashtags, headlines, and comments. That made sense at first, but these days, it’s just a small piece of the puzzle. More and more, the conversations that matter most are happening out loud, in visuals, or in on-screen text that never makes it into a caption.

If your tools only read captions, you’re missing out on quantity, and, you’re missing out on the real, unfiltered thoughts people share out loud or show on screen, instead of carefully typing and editing.

What Changes Depending on the Source

Not every mention is created equal. If you treat them all the same, you risk missing the important stuff or letting the real gems get lost in the noise.

Written and captioned mentions are the easiest to spot and double-check, but they’re also the most polished. People edit captions all the time. But what someone says out loud, mid-sentence, in a 40-minute video? That’s usually straight from the heart.

Spoken mentions, picked up by Automatic Speech Recognition (ASR), are usually more honest and off the cuff. Think of a creator sharing their real thoughts in a product review, a podcast host chatting about the latest news, or someone diving deep in a YouTube essay. This is where people say things they might never write down and where you’ll often catch the first hint of a disclosure, a bold claim, or even a brand-damaging comment.

Visual and on-screen text mentions, found using Optical Character Recognition (OCR) and Computer Vision, are a whole different story. Maybe your logo pops up in the background of a video, or someone shares a screenshot of a text conversation, or there’s a handwritten sign at a protest. None of this appears in a caption, and it’s not “written” in the usual way, but it can be the clearest proof of a brand connection you’ll ever see.

Here’s what the numbers show: about 75% of brand mentions in video content live only in the spoken words, not in the title, description, or hashtags. On Instagram, 66% of brand mentions can only be found with OCR or AI-generated image captions. That means two out of three things people say about your brand on Instagram are invisible if your tool only reads text fields.

The Three Technologies Behind Real Multimodal Filtering

To really get the full picture, you need three types of AI working together on the same content, all at once, not in separate corners.

  • ASR (Automatic Speech Recognition) transcribes spoken content across podcasts, YouTube videos, TikTok voiceovers, and livestreams, making every spoken word searchable in the same way a caption is.
  • OCR (Optical Character Recognition) reads text that appears visually within a frame, screenshots, signage, on-screen graphics, embedded captions, and content that exists only as pixels until OCR turns it into searchable text.
  • Computer Vision identifies logos, products, faces, and visual context even when nothing is written or said at all.

Pendulum brings all three together in one Snippet System. That way, a spoken phrase from a podcast, a logo spotted in a video, and text from a screenshot all show up in one place no more jumping between different dashboards.

Why Filtering by Demographic and Geography Matters Once You Can See Everything

Multimodal detection helps you see what’s out there. But the next step is figuring out who’s saying it, and where. Once you can spot every kind of mention, filtering by geography and demographics turns "we found the conversation" into "we know who’s behind it and why it matters."

This difference really matters when you’re trying to separate real insights from background noise, or decide if an alert needs your attention. A spoken mention from a small, local group is a whole different story than the same words shared by a big account with national reach. If your system only looks for keywords, it can’t tell the difference.

Source-Aware Filtering Is What Makes Monitoring Agentic

This is where filtering by source type ties into the bigger move from static rules to agentic monitoring. A rules-based system can flag a keyword, sure. But it can’t tell the difference between a spontaneous comment in a podcast and a carefully crafted caption, or spot that a protest sign caught by OCR is a bigger deal than the same words in a tweet. Real judgment comes from understanding both the source and the context; this is what agentic AI really means for brand intelligence.

Treating every mention like it’s the same isn’t just a small oversight. It’s why so many monitoring systems either flood teams with unhelpful alerts or miss the one mention that really matters just because it was spoken, not typed.

See How Your Social Listening Measures up With Our Free Social Listening Audit

Multimodal intelligence · Pendulum

Frequently Asked Questions

What is social media audio monitoring?

Social media audio monitoring uses Automatic Speech Recognition (ASR) to transcribe spoken content across podcasts, videos, and livestreams, making brand mentions said out loud just as searchable as a written caption or hashtag.

Why does the source of a brand mention matter?

The source changes what a mention means. Written captions are curated and easy to fact-check. Spoken mentions, captured via ASR, tend to be more candid and unscripted. Visual mentions, captured via OCR and Computer Vision, often carry no caption at all. Treating all three the same buries real signal in noise.

What is OCR social listening?

OCR (Optical Character Recognition) social listening reads text that appears visually within images and video frames — screenshots, signage, on-screen graphics — and converts it into searchable text, surfacing brand mentions that never appear in a caption or hashtag.

How much brand conversation happens outside of text captions?

Roughly 75% of brand mentions in video content exist only in the spoken transcript, and 66% of Instagram brand mentions are detectable only through OCR analysis or AI-generated image captions rather than through written text fields.

What is multimodal brand intelligence?

Multimodal brand intelligence analyzes text, audio, video, and image content together using ASR, OCR, and Computer Vision simultaneously, rather than defaulting to text-only keyword tracking, so that spoken, written, and visual mentions are all captured and understood in context.

What does demographic brand filtering add to monitoring?

Demographic and geographic filtering shows who is driving a conversation and where, once multimodal detection has already surfaced it — turning "we found the mention" into "we understand who's amplifying it and how much reach it actually has."

Times Have Changed. So Should Your Tech Stack

See how brands are upgrading their strategies with Pendulum Social Intelligence.