Dear Marketing Director,
I’m writing this after another long stretch of audits, and I’ll be blunt with you the way I’ve been blunt with clients lately: the “collect everything” mentality has become one of the most expensive habits in enterprise marketing.
For years, agencies and platform sales teams sold executives on the idea that more data automatically meant more intelligence. Track every click. Save every parameter. Keep every old event stream in BigQuery. Hold onto everything because maybe, someday, someone might want to analyze it.
That advice has aged badly.
In 2026, “just in case” data is not a strategic asset by default. In many organizations, it is a liability sitting quietly on the books, waiting for an audit, a breach, a legal review, or a ballooning cloud invoice to make the problem impossible to ignore.
I’ve been inside enough messy environments now to tell you this isn’t theoretical. I keep seeing the same pattern across higher education, government, and B2B organizations with large, complicated websites. Someone added tags over the years with good intentions. Then more vendors were layered in. Then a CRM sync got bolted on. Then an ad platform connector. Then a form tool started passing values nobody meant to capture. A few years later, nobody can clearly answer a simple executive question: what are we collecting, why are we collecting it, where is it going, and who is responsible for it?
That’s the real issue. If nobody can explain the business purpose of the data in plain English, you probably should not be collecting it.
I was in one audit recently where a client believed they had a fairly conservative analytics setup. On paper, it looked normal. GA4. Tag Manager. A few ad pixels. BigQuery export. Nothing dramatic. But when I traced the flow end to end, we found URL parameters carrying email addresses from campaign links, internal search patterns being logged far longer than anyone realized, and form interactions that could expose values no one in leadership thought were leaving the browser. Nobody had malicious intent. It was just years of accumulation, handoffs, and assumptions.
That’s the part that should concern executive leaders. Most data problems in marketing are not caused by one catastrophic decision. They’re caused by dozens of small “sure, keep it” decisions that nobody revisits.
And yes, the cost is financial in the obvious sense. BigQuery costs add up when you export everything, retain too much, and let event design sprawl. I’ve seen organizations pay to warehouse rows nobody uses, nobody trusts, and nobody maps back to a business KPI. That’s bad enough.
But the larger cost is operational and legal.
Every unnecessary identifier, every stray query parameter, every over-retained event record expands your attack surface. Every extra row becomes one more thing security, legal, compliance, and IT may have to explain later. In a breach scenario, your cleanup burden is tied not just to whether data was exposed, but to how much unnecessary data you allowed to exist in the first place.
That’s why I’ve changed the way I talk about analytics architecture with clients. I’m spending less time arguing about tools and more time arguing for discipline.
I still believe in measurement. I still believe in attribution where it’s useful. I still believe marketing leaders need strong analytics to defend budgets, improve lead quality, and understand what actually drives outcomes. But I no longer accept the lazy assumption that the answer is to gather every possible signal and sort it out later.
Later never comes.
What I’m recommending now is a much tighter model built around data minimization, deterministic business signals, and controlled collection at the server boundary.
In plain English, that means we stop treating the browser like an open faucet feeding every platform downstream. We get much more intentional about what leaves the page, what gets transformed, and what gets stored.
When I review Google Tag Manager containers now, I’m looking less for cleverness and more for restraint. I want to know whether the events being pushed into the data layer correspond to actual business decisions. Are you measuring application starts, qualified lead submissions, service completions, donation transactions, or other concrete outcomes? Good. That’s useful. Are you collecting a bunch of hover behavior, redundant scroll events, or verbose interaction noise that nobody uses except to create dashboard theater? That’s usually where I start cutting.
I’ve had some uncomfortable conversations with teams about this, especially when someone has spent years building a very “rich” tracking setup. But richness is not the same as clarity. A haystack is not a strategy.
One of the most important shifts I’m recommending is server-side masking. If you’re in a complex B2B, government, or higher ed environment, I increasingly see this as less of a “nice to have” and more of a practical control point.
When data goes server-side first, you gain the opportunity to inspect it before it gets handed to GA4, Meta, LinkedIn, or anyone else. That matters. It gives you a chance to strip sensitive query parameters, redact accidental PII, normalize payloads, and hash identifiers where a persistent match is necessary for limited use cases. I’m talking about controls like removing email, phone, or zip_code from incoming URLs before they ever hit your warehouse, or applying SHA-256 hashing at the server layer so the raw value never becomes part of your reporting exhaust.
I want to be clear here: hashing is not magic and it is not a loophole. It is one control within a larger governance system. If your upstream collection model is reckless, no amount of hashing is going to save you from bad architecture.
The same goes for IP handling. Vendors will tell you they anonymize, minimize, or process responsibly, and sometimes they do. But I’d rather put the strongest possible control at the boundary I manage than rely on downstream assurances I don’t control. That’s the philosophy I keep returning to with clients: own the part you can govern.
I’m also telling teams to get much more aggressive about retention and deletion. If a tag hasn’t informed a report, decision, or optimization in months, why is it still there? If a field cannot be mapped to a KPI, audience strategy, or operational need, why are you paying to store it? If an old connector exists because a vendor set it up three years ago and nobody wants to touch it, that is not governance. That’s procrastination with cloud costs attached.
A few weeks ago, I was reviewing an environment where leadership was worried about performance, privacy, and reporting quality all at once. They assumed these were separate problems. They weren’t. The same undisciplined collection model was making the site heavier, the reports noisier, and the compliance posture weaker. Once we started pruning tags, reducing unnecessary client-side calls, and tightening the event model around real conversion points, the picture got clearer fast. Less noise. Lower storage waste. Better confidence in what the dashboards actually meant.
That’s another point I wish more executive teams heard: data minimization is not the enemy of insight. Very often, it is the precondition for insight.
When you only collect what matters, reporting becomes more human-readable. Your dashboards stop looking like a junk drawer. Your team can explain performance without drowning everyone in jargon. And when somebody from legal, procurement, or security asks what your measurement system is doing, you can answer without sounding like you’re reading from a vendor brochure.
If you lead marketing in a large organization, my advice to you is simple even if the implementation is not. Push your team to justify collection in business terms. Ask where PII could leak accidentally. Ask what reaches BigQuery and why. Ask whether your current setup is preserving meaningful signals or just preserving history because nobody wants to delete anything. Ask whether your analytics architecture reflects your organization’s risk tolerance, not just your agency’s appetite for more data.
I know this can feel like one more governance burden in an already complicated stack. But from what I’m seeing in the field, this is no longer optional cleanup work. It’s foundational. The old habit of collecting first and thinking later is colliding with privacy expectations, legal scrutiny, infrastructure costs, and basic operational sanity.
Your data should help you make decisions, not create a larger mess for your security team, your legal team, and your finance team to clean up later.
That’s where I’ve landed after these audits. Measure what matters. Redact aggressively. Mask at the server boundary. Keep the signals that serve the business. Delete the rest.
Sincerely,
Marcus Sanford

