Pages that have appeared in your results
Your domain is showing up for words you never wrote: a product you don't sell, another language, another line of business. Those pages exist, they carry your domain name, and someone else wrote them. What follows is the order in which the steps should be taken — measure the actual scope, remove, then bring the index back down — and why that order matters.
Visible scope and actual scope
A "site:" search followed by your domain shows what a search engine has retained about you. It's the right starting point, and it's also the trap in this situation: that list tells you what the crawler has explored and kept, as of a given date. It says little about what your server actually contains. The pages it passed over are served exactly like the others, and they're waiting their turn to be crawled.
Two practical consequences. First: the number you see is a floor, it tends to rise in the days that follow, and counting results amounts to measuring the crawler's appetite rather than the extent of the work. That extent is measured in the files, on the server.
Second: these pages sometimes reveal themselves to the crawler while hiding from you. The same URL serves foreign text to the search engine and your usual page to your browser. So a URL that looks clean from your own machine may be exactly the one showing up in the results — hence this page's guiding principle: trust what the server contains, verify what the index displays.
The order of operations
Each step protects the next. Hiding before you've removed buys you a reprieve; removing before you've copied loses the date; and removing everything without looking for what's writing it means doing this entire job over again later.
Ask each search engine what it has kept about your domain
The "site:" search followed by your domain, with no space, lists the URLs a search engine has kept from you. Run it on each of the ones that send you visitors, and note what recurs: the same folder prefix, the same URL pattern, the same language, the same topic. That pattern is worth more than the list itself — it's what you'll go looking for on the server, in step 03.
Open index coverage in Search Console
Google Search Console gives you your site's index coverage: the URLs it knows about, those that are indexed, those that were excluded and the reason for each exclusion. It sees further than a "site:" search, and it can be exported. Date this list and keep it: it's your baseline for what comes next, and the document that will later show the index coming back down.
Look for the same pattern on the server itself
This is where the true extent becomes visible, and only here. Before touching anything at all, copy the files AND the database in their current state onto separate storage: once removed, what has been removed becomes hard to date. Then, over FTP or SSH, sort by modification date and work back to your own last deployment. Look at the upload folders, the caches, the temporary directories — the places where your site writes on its own.
Remove the pages, then make their URLs respond with a 404 or 410
Removing the file is the action; the server's response is the proof. A removed page must respond with a 404 or a 410 to leave the index for good. Check this URL by URL on a sample of each pattern identified in step 01: a URL that still returns a 200, even on a blank page, remains a page in a search engine's eyes. While you are at it, review what an .htaccess file or the server configuration adds — a rewrite rule can replay URLs whose files are gone.
Request temporary removal, knowing what it does
Search Console's temporary removal tool hides a page from results for the time it takes the search engine to recrawl it. It is a curtain, and it is useful: when foreign pages are being seen by your customers, stopping them from showing right away has value. It is, on the other hand, a poor substitute for removing the page itself — if the URL still responds, it comes back when the period ends. The order fits in one sentence: remove first, hide second, while the index catches up.
Establish the exact list of files that were added
This is the point where the human eye reaches its limit, and the one that decides whether all of this starts over. A live site carries tens of thousands of files, and the one that was added looks just like its neighbours. The question that can actually be answered, however, is the reverse one: how many files can justify their presence, and which ones cannot? That is what our analysis does — it compares your site with the authentic code published by the vendors, names every discrepancy, and delivers the whole thing signed.
Pages come back as long as the factory that makes them stays in place
These pages have a source. Most often, they are files dropped onto the server, somewhere your site had write permission. Sometimes a single added component manufactures them on demand, from a list of words: their number then becomes a question without an answer, since it will write as many as the bot requests.
That is what makes removal alone disappointing. You delete the addresses identified in step 01, the index goes back down, and new addresses appear — different words, different folders, the same domain. As long as the component that writes them remains in place, the work is redone identically.
There is one more reason to trace it back that far: whatever wrote those pages got in through somewhere, and a door that has been used once will be used twice. Removal deals with what is visible; the exact list of added files deals with what produces it.
Step 06, done for you, with the means to prove it
The analysis answers the question of step 06 by turning it around: instead of looking for what is known to be bad, it asks every file to justify its presence. A file that matches the authentic code published by its publisher is accounted for. The others are named, one by one, with their date and their location. That is exactly the list that is missing when you don't know where to start with the server.
You receive the exact number, the list of files, the probable date of entry, and three signed documents: the report, the certificate of the state obtained, and the document you can hand to your host, your client or your insurer. Each one can be verified with us, by anyone, free of charge and forever — that is what makes them enforceable. Record, €99 per site; with the restoration of the publishers' files, €199.
There is no free analysis here, and that is deliberate. A free analysis would almost always answer "nothing found" — the expected good news, given without proof, to people who didn't need it. We sell certainty, rather than worry. Verifying a document, on the other hand, is free for everyone and forever: it is the half that must remain open for the other to be worth anything.
One caveat, written here as it is on every document: a fully accounted-for site is a site in which every file has accounted for itself — which is already a great deal, and is still not the same thing as a secure site. A weak password, or an up-to-date but vulnerable extension, remains beyond the reach of a file comparison.
What holds up after removal
Once the pages have been removed and their URLs return 404 or 410, the index drops back as the crawler makes its passes. That pace is out of your hands: it's the only place on this page where patience takes the place of an action. Track it in the indexing coverage report dated from action 02, and keep the record stating what you removed and when — that's the document that answers the questions raised afterwards.
What really changes what comes next boils down to one thing: knowing the same day that an executable file has changed on your server, rather than finding out by discovering your domain on foreign-language keywords. Monitoring states it precisely — when the attestation breaks, it announces that this site differs from the one that was attested. Your own deployment breaks it in the same way: that's why we re-attest after every deployment, and why a break without a deployment warrants a close look. Watch, €99 per site per year, in a single payment.
How many files on your site can justify their presence?
You'll have the exact number, each file named, the likely date, and three signed documents that someone else can verify.