Back to blog
Audit

How to run a sales audit on your Instagram and WhatsApp DMs

5 min read
How to run a sales audit on your Instagram and WhatsApp DMs

A social media sales audit is a one-time review of a closed window of past Instagram, Facebook, and WhatsApp conversations, graded one thread at a time against the same short list of yes or no columns, then totaled. It is Sales QA run by hand: a spreadsheet, an afternoon, and one rule applied identically to every row.

It is not the daily pass, which asks who is still waiting right now. An audit is retrospective: it names the step that broke repeatedly last month, because every conversation in it has ended. For what a review grades inside one thread, see Sales QA for social conversations.

Grade the threads, do not read them. An open question invites a defense, she went quiet, it was our busy week, and a defense does not total. A column of no's does.

What window do you pull, and what counts as one row?

The last 30 full days, ending yesterday, and one row per customer thread on one channel, however many messages it holds. Thirty days is long enough for a repeated mistake to appear twice, short enough that people still remember the week. Grading message by message counts a chatty thread several times and a one-line price question once.

Work one channel at a time and exclude nothing, not spam, not fan mail, not the thread you know went badly. WhatsApp means scrolling the phone the business number lives on, the slow part people skip, and you can only review the history that device still holds. No software fetches last month's WhatsApp threads for you either: monitoring starts at connection. None of the three exports a usable list, so type a row per thread as you read it.

Too long for one sitting? Write the sampling rule down before you read a single thread, because a sample chosen afterward is a sample of the threads you already remember: every third conversation in date order, leads and non-leads alike. An older thread counts if the customer wrote inside the window, a comment that moved into a DM is one row dated at the comment, and one person on two channels is two rows.

Which columns belong on the audit scorecard?

Ten, and every one of them adds up at the bottom of the sheet. Write them before you open the first thread: a column invented at row forty cannot be applied to the thirty-nine above it.

ColumnWhat goes in the cellWhy it earns a column
LabelFirst name, date, channelPoints back to the evidence
ArrivedClock time of their first messageSplits working from after hours
LeadYes or noThe denominator for every line below
IntentPricing, purchase, booking, availability, trust, objectionWhat a good reply owed them
First business replyHuman, auto-reply only, or noneAn acknowledgment looks like an answer
Reply timeTo that human reply, blank if noneBlank is the costliest cell here
Answered in fullYes only if they did not ask againEvery re-ask is a place to leave
Next stepYes only if your last message names a time, names an action, or asks a one-word questionWhere replied and booked separate
Last message byUs or themTheir turn means dropped, not lost
OutcomeBooked, bought, declined, dropped, unknownUnknown counts results living in memory

Fill the intent cell only where Lead says yes, using the six buying intents defined in how to sort social conversations by buying intent. Never score a thread out of five: a rating does not total, and two people never agree on a 3. The first-reply column matters because an unanswered row is not neutral. In the Sprout Social Index, 2025, 73% of social users said they will buy from a competitor when a brand does not respond to them on social.

How do you grade a thread without arguing about it?

Read the customer's last message first, then the thread from the top, then fill the row and move on. Grade what is written, not how the team did.

  1. Grade the screen, not the intention. Meaning to circle back leaves no trace, so it scores a no.
  2. Never grade tone. Friendly and terminal are not opposites, and warmth does not total.
  3. Grade before you see the outcome. A booking makes a sloppy thread read as fine, a loss makes a clean one read as guilty.
  4. If two people would disagree on a cell, the column is worded wrong. Fix the wording, regrade, move on.

That last rule is why the next-step column names three concrete things. Did we move them forward is a debate. Does your last message name a time, name an action, or ask a one-word question is a lookup. Three cells flatter whoever fills them: an acknowledgment counted as a reply, a redirect counted as an answer, an outcome typed from memory.

One graded thread that nobody at the clinic would have flagged

A parent asks, gets a true answer, says thanks. In an inbox that is a good exchange. In the sheet it fails one cell, and it is the cell somebody will contest.

Sun 20:41, Owen Marsh, WhatsApp: My son is 7 and nervous about dentists. Do you see kids that age?

Mon 11:26, clinic: Yes, we see children from age 4. A first visit is just a look around, no treatment.

Mon 11:31, Owen: Good to know, thanks.

The row, from a composed illustration at a fictional clinic: Owen M Aug 9 WhatsApp, 20:41 after hours, lead yes, trust intent, first reply human, 14h 45m, answered in full yes, next step no, last message by them, outcome unknown.

Somebody will contest that cell. We answered honestly and he thanked us, so what more was there to say? By the column's wording: no time named, no action named, no one-word question asked. One row like this is unremarkable. It becomes information when thirty rows sit under it and the next-step column is empty down the page.

How do you total the sheet and decide what to fix first?

Total it as a funnel, each line a subset of the one above: conversations pulled, leads, leads that got a human reply, leads whose questions were fully answered, leads whose last message carried a next step, and leads that reached a booking, a payment, or a call. If a line exceeds the line above it, some column is loose enough to be counted two ways.

Then take the largest single drop between two consecutive lines, discounting any drop below an earlier break, because every line beneath a break is measured on a smaller population than it deserves. When two gaps are close, repair the later one: it sits on warmer buyers.

What the totals showThe broken stepThe first change to make
Biggest drop: conversations to human replyAnswers follow the message list, not the buyersSort by intent before answering
Lead yes, first reply auto-reply onlyThe acknowledgment counts as the answerWork that list as unanswered
After-hours reply times dwarf your working-hours onesThe evening has no ownerName one owner, compare again next month
Biggest drop: human reply to answered in fullStandard answers cover half the questionOne complete answer per intent
Answered in full high, next step falls awayThe template ends at the answerEnd answers with a time, an action, or a question
Every column looks fine, outcome mostly unknownThe sample, not the processRe-pull by the written rule: you graded what you remember

The example Sales QA report on the Reploom homepage shows the finished shape. One sample week: 420 conversations checked, 96 potential buyers, 74% replied or handled, 9 still waiting, 41m average reply time inside working hours against 1h 23m outside. Sample figures from one example account, not research and not a benchmark. Note how it ends: 7 sales opportunities stalled this week. Six numbers get filed. One sentence gets assigned.

How do you re-run the audit and prove the fix worked?

Same columns, same wording, same window length, a new 30 days, and compare only the lines you meant to move. Never redefine a column: if the wording was loose, add a new one rather than editing the old one, which silently rewrites the first result.

Compare fractions rather than raw counts, because counts move with demand: a busy month lifts conversations, leads, and bookings together while every broken step stays as broken. Compare against your own first run, not a published figure: no outside number says what a healthy next-step rate is for a business your size. The fix worked if the gap you targeted narrowed, and it did not fail if that gap moved one line down the funnel.

When does a hand-run audit stop working?

At volume and at frequency. It covers one window once, not three channels every week, and it is graded by the person whose process is being graded. It also ages: the leak it names is the leak you had. By hand, quarterly is realistic. Weekly is not.

Reploom is AI Sales QA for social conversations. It reviews Instagram, Facebook, and WhatsApp messages, comments, and mentions after they happen, scores each one, and groups repeated misses into named patterns in a weekly report with the source conversations as evidence. It does not draft, send, or approve replies, it is not an inbox and not a chatbot, and it does not replace the people answering DMs. On WhatsApp, monitoring starts at connection, so how far back a review reaches depends on available history.

Run the manual version first. It costs an afternoon and names which of the ten columns is your problem. To see those columns filled in every week, by something that is not grading its own work, book a demo.

See which leads are getting lost

Reploom reviews social sales conversations and turns repeated lost-lead patterns into a weekly Sales QA report with evidence and fixes. See how Sales QA for social conversations works.

Book a demo