Tech
I Built a Smarter Privacy Shield for Your Browser. Then Discovered Why Dumb Noise Wins.
I Built a Smarter Privacy Shield for Your Browser. Then Discovered Why Dumb Noise Wins.
And why the result is more interesting than if I’d just won.
Your browser is leaking. Not through cookies — those are easy to block. Through something subtler: the unique fingerprint of how your device renders fonts, handles WebGL, reports its screen resolution, and dozens of other signals. Combine enough of them and you’re identifiable across the entire web, even in private mode, even with a VPN.
The standard defense is noise injection — add small random perturbations to these signals so trackers can’t pin you down. Simple idea. But there’s a question nobody had answered cleanly: does it matter where you put the noise?
Intuitively, it should. Some browser attributes matter enormously for identification (canvas fingerprints, WebGL hashes). Others barely matter at all (CPU core count, memory). If you’re working with a limited noise budget, concentrating it on high-impact features should give you more privacy per unit of disruption.
That’s the idea behind CANO: Context-Aware Noise Optimization. I built it, ran 68,885 experiments across 12 datasets and 6 different attack strategies, and found something that surprised me.
How CANO Works
CANO uses a Random Forest to measure each feature’s contribution to re-identification (permutation importance). Then it allocates noise in proportion to that importance — high-impact features get more noise, low-impact features get less.
The key formula is deceptively simple:
δᵢ = ε · (wᵢ · n_features) · sign(zᵢ)
The n_features scaling factor is the critical piece. Without it, distributing a budget across n features means each feature gets 1/n of what a uniform strategy would apply — you’d be using n times less noise overall. That scaling correction alone was the single biggest performance improvement in the whole project.
I also added a minimum weight floor (w_min = 0.1) — no feature ever gets truly zero noise. The reason: an adaptive attacker will notice which features are unprotected and simply use those instead.
On top of the static allocation, I trained a Deep Q-Network (DQN) through adversarial co-evolution — alternating between the defense protecting data and the attacker retraining on protected outputs. Thirty rounds of this arms race.
The Part Where I Lost
Against a known attacker — one trained specifically on your noise strategy — Gaussian noise wins. Not by a little:
Strategy
Accuracy Reduction
Gaussian
0.395
Laplace
0.239
FGSM
0.212
PGD
0.129
CANO
0.112
C&W
0.001
CANO ranks fifth out of six. Uniform random noise beats a carefully engineered feature-importance system by 3.5×. That’s a humbling result.
Even more humbling: the DQN policy, after 30 rounds of adversarial training, converged to a Gini coefficient of 0.009 — essentially uniform allocation. The reinforcement learner, given complete freedom to optimize noise distribution, independently rediscovered Gaussian noise. The game-theoretic equilibrium against an adaptive attacker is uniform noise.
So why did I spend months building this?
The Part That Makes It Interesting
Here’s the thing about that Gaussian result: the attacker knew exactly which strategy was being used.
Real-world privacy systems don’t get that luxury. A fingerprinter doesn’t know whether you’re running CANO, Brave’s randomization, a custom VPN-level perturbation, or something else. They have to generalize.
So I measured something different: transfer efficiency — how well accuracy reduction against one attack model transfers to a different attack model the defender never trained against.
Strategy
Adaptive
Transfer
Ratio
CANO
0.112
0.271
2.41×
FGSM
0.212
0.474
2.23×
Laplace
0.239
0.392
1.64×
PGD
0.129
0.145
1.13×
Gaussian
0.395
0.410
1.04×
C&W
0.001
−0.016
n/a
Gaussian’s transfer ratio is 1.04×. Its entire adaptive-attack advantage essentially disappears when the attacker uses a different model. CANO’s transfer ratio is 2.41× — its perturbations are more than twice as effective against unseen attack models as against the known one.
The mechanism makes sense in retrospect. CANO targets the attributes that structurally matter for re-identification — the ones any classifier would have to rely on, regardless of architecture. Gaussian noise is indiscriminate, and its effectiveness against a specific classifier doesn’t generalize because that classifier’s decision boundaries are specific.
What Real Data Showed
I evaluated against the FP-Stalker corpus — 776 real users, 13,674 actual browser fingerprints, 34 real attributes collected over weeks. This is the closest thing to the real deployment scenario.
On synthetic data, CANO’s gap to Gaussian is dramatic (e.g., 0.003 vs 0.514 on one dataset). On real fingerprint data:
CANO 0.276 vs Gaussian 0.340.
That’s a gap, but it’s a much smaller one. The synthetic datasets have artificial importance concentration — a few features dominate completely. Real browser fingerprints have more redundancy across attributes, which is exactly the scenario where importance-weighted allocation pays off.
The Noise Quality Difference
CANO and Gaussian apply similar amounts of noise (L2 = 0.435 vs 0.595), but their noise structure is different:
Strategy
SNR (dB)
Sensitivity
Gaussian
9.7
+0.110
CANO
15.5
−0.226
CANO’s SNR is 15.5 dB vs Gaussian’s 9.7 dB — the noise is less perceptible to legitimate services while still disrupting identification. The negative sensitivity signature reflects that CANO concentrates on importance-ranked features rather than high-variance ones.
For a privacy tool that real users would actually deploy, that SNR difference matters. Less noise disruption to legitimate web services means fewer broken pages, fewer CAPTCHA prompts, fewer sites that detect something is wrong.
What I’d Do Next
The two most interesting directions:
Formal guarantees. CANO’s allocation is heuristic — it has no differential privacy guarantees. Deriving a formal ε-DP bound for the allocation mechanism would make it deployable in high-stakes contexts.
Online RL. The current DQN trains offline. A policy that updates in deployment — adapting as it observes which strategies the live fingerprinter is using — would be substantially more powerful.
More real data. One real-world dataset (FP-Stalker) is better than none, but the HTillmann and BrFAST extended datasets would strengthen external validity considerably.
The Honest Summary
CANO loses on the metric that’s easiest to measure. It wins on the metric that actually reflects deployment reality.
Uniform noise is the optimal defense when the attacker knows your strategy. Feature-importance allocation is better when they don’t — and they usually don’t.
The RL equilibrium result is the finding I keep thinking about: given complete freedom, an optimizer independently discovers that against adaptive adversaries, you should distribute noise uniformly. That’s not a failure of the approach. It’s a proof that the approach was asking the right question.
The full paper (with all 68,885 experiment configs, 7 figures, and complete methodology) is available on [Zenodo / HuggingFace Papers]. Code at [GitHub link].
All analysis was run on real browser fingerprint data from the FP-Stalker corpus (Vastel et al., IEEE S&P 2018) and three synthetic benchmark datasets.