Publish a position paper in French, Dutch and English and you have not created one document — you have created three that a crawler has to fetch, compare and decide between. This article is about what that costs, how to tell whether you are paying it, and what the Indexing Hub can and cannot do about it.

There is a particular kind of frustration that belongs to organisations who publish a lot and are read by few: the document exists, the URL works, colleagues can open it, and yet a search for its exact title returns nothing. Nobody has removed it. It simply was never taken into the index.

For a federation, an advocacy office or a consultancy that publishes consultation responses, briefings and position papers on a schedule, this is not a cosmetic problem. A document that is not indexed cannot be cited, cannot be found by the official verifying your track record, and cannot be linked to by anyone who did not already have the address. The chapters below set out where discovery actually breaks, and what the Semalt indexing tools change about it.

Mechanics · Four stages, four failure points

Discovery, crawling, rendering, indexing — and why only the last one counts

The path from publishing to appearing in results has four distinct stages, and they fail differently. Discovery is the search engine learning that an address exists. Crawling is fetching it. Rendering is executing whatever the page needs in order to have content. Indexing is the decision to store it and make it retrievable.

Most conversations about this collapse the four into one word. That is why they go nowhere: "the page isn't indexed" describes an outcome, not a cause, and each of the four stages needs a different intervention.

StageWhat has gone wrongWhere you see itWhat actually helps
DiscoveryNothing points at the addressURL absent from every reportSitemap entry, internal link, submission
CrawlingFetched late, rarely, or not at allLong gaps between bot visitsFewer competing URLs, faster responses
RenderingContent only exists after script executionEmpty fetched sourceServer-side output of the main text
IndexingCrawled, then declined"Crawled — currently not indexed"A better document, or fewer near-copies
Submitting a URL is not the same as being indexed. A submission is a request to look. It puts the address in a queue and nothing more. If the page is thin, near-identical to another, or contradicted by its own canonical tag, it will be fetched and then declined — and it will be declined faster the second time. No tool changes this, and any tool that promises otherwise is describing something it does not control.

That distinction is the whole reason indexing work is frequently disappointing. Teams submit, see the counter rise, and conclude the job is done. The counter is measuring submissions. Presence in results is measured somewhere else entirely, and the gap between the two numbers is the actual state of your site.

Scale · Three versions, one budget

What a trilingual publication schedule does to the arithmetic

Brussels organisations frequently run three language versions: French, Dutch and English. It is rarely a vanity decision — the domestic audience splits along a language line, and the international membership reads the English version because it is the only one they all share. The editorial logic is sound. The crawl arithmetic is less kind.

Every published document becomes three addresses. Every tag page, every archive page, every author page and every pagination step multiplies by three as well. A site with 400 documents in one language is a small site; the same site trilingual, with archives and filters, is routinely twenty thousand crawlable addresses. Nothing about that is anyone's fault, and all of it consumes the same finite attention.

×3
addresses per document
×3
archive and tag surfaces
1
crawl budget to share
1,000
URLs per day in the tracker

Crawl budget is not a quota you can top up. It is the practical outcome of how much a search engine considers your site worth fetching, adjusted for how quickly your server answers. You influence it in exactly two ways: by making the site faster, and by reducing the number of addresses competing for the same attention.

The uncomfortable implication. If a crawler spends its visits on filtered archive pages and session-parameter duplicates, the position paper you published yesterday waits. Discovery problems are usually not about the new document at all — they are about everything else on the site that got there first.
Duplication · The surface nobody counted

Three languages are not three duplicates — unless you make them so

Language versions are legitimate. A properly declared French, Dutch and English set is not duplicate content, and search engines handle it well when the declarations are consistent. The problems come from the machinery around the versions rather than the versions themselves.

Common fault

Partial translation fallback

The English URL exists for every document, but where no translation was made it serves the French text under an English address.

  • Creates genuine duplicates
  • Breaks the language declaration
Common fault

Asymmetric declarations

The French page names the English one as its alternate, but the English page does not name the French one back.

  • Unreciprocated pairs are ignored
  • Check both directions, not one
Common fault

Canonical pointing across languages

Every version canonicalises to the French original, which asks the search engine to drop the other two entirely.

  • Each version canonicalises to itself
  • Language pairing is a separate declaration
Common fault

Filter and parameter explosion

Archive pages accept a topic, a year and a language as parameters, and each combination becomes a crawlable address.

  • Multiplies by three like everything else
  • Almost never worth indexing

The first of those deserves particular attention because it is so easy to build by accident. A publishing system that never returns a 404 for a missing translation feels helpful. What it actually does is create a large set of addresses whose content is identical to another address in a different language directory, and near-identical content at scale is the most reliable way to spend a crawl budget on nothing.

There is a second version of the same trap that has nothing to do with language: near-identical documents in the same language. Twelve consultation responses that share a boilerplate introduction, a standard methodology section and a signature block may differ in only three paragraphs out of forty. Each is a legitimate document. From a crawler's perspective they are close enough that some of them will be crawled and then declined.

A test that takes five minutes. Take one of your less prominent documents, search for a distinctive twelve-word sentence from the middle of it in quotation marks, and see whether the document comes back. If it does not, the problem is indexing, not ranking, and no amount of keyword work will address it.
Tooling · What the hub does

The Indexing Hub in practice

The indexing area of the dashboard exists to do the mechanical part properly: get addresses in front of the crawlers, and then record what happened to each one. It does not replace the structural work above, and it is most useful once that work is done.

Indexing Hub · Submission

Getting addresses into the queue

Individual URLs, bulk batches and sitemaps, submitted through the IndexNow API to GoogleBot and BingBot.

included · per account
  • A daily ceiling that is real. The URL tracker works to a budget of 1,000 URLs per day per account, which is a design decision rather than a limitation to work around.
  • Bulk where bulk makes sense. Up to 10,000 URLs in a single batch, for a migration or a first import rather than for weekly publishing.
  • Sitemaps by file or address. Upload the file or point at its URL; nested indexes are parsed recursively to three levels, up to 1,000 sitemaps in one job.
  • Queue discipline. Two sitemap jobs run at once and up to twenty wait in line, so a large import does not block a routine submission.
1,000
URLs per day
10,000
URLs per batch
3
levels of sitemap nesting
2 / 20
jobs running / queued
Indexing Hub · Observation

Recording what the bots actually did

A per-URL log of visits, statuses and errors, with live counters for submitted, found and failed addresses.

included · continuous
  • Timestamped visits. Each URL carries a record of when a bot came, what it received and what went wrong, which turns an argument into a lookup.
  • Three counters, not one. Submitted, found and failed are separate figures, and the distance between the first and the second is the number that matters.
  • Error detail per address. A 500 during a crawl window and a redirect chain look nothing alike in the log, though both produce the same absence in results.
  • Filterable by site tag. Where several properties sit in one account, tags act as a single filter across the whole workspace.

The observation half is the more valuable of the two, and it is the half people ignore. Submission is a button. The log is evidence — the thing that lets you say, with a date attached, that a document was fetched on the eleventh and returned a server error, rather than speculating about whether Google "likes" the page.

Diagnosis · Reading the gap

Submitted, found, failed: what the three counters tell you

Three numbers, read together, describe your situation more accurately than any single indexing score could. Read separately they mislead, which is why the counters are presented side by side.

submitted
addresses put in the queue
found
addresses a bot actually fetched
failed
fetches that returned an error
  • Submitted high, found low. Crawlers are not reaching the addresses, or are reaching them and getting something other than a page. Look at server responses and redirect chains before touching the content.
  • Found high, still absent from results. The pages are being fetched and declined. This is a content and duplication problem, and resubmitting will not alter the decision.
  • Failures clustered in one directory. Almost always a template or a server rule rather than the individual documents. Check one page and you have checked all of them.
  • Failures spread evenly across the site. Points at infrastructure — response times under load, or a rate limit that treats crawlers as unwanted traffic.

The second pattern is the one that costs organisations the most time, because it feels like a technical problem and is not. When a document is fetched and then not stored, the search engine has looked at it and judged it not worth keeping — usually because something extremely similar is already indexed, or because the page has very little on it beyond a title and a download link.

The PDF question, since it always comes up. A briefing published only as a PDF can be indexed, but it competes badly and it strips out your internal linking. The reliable pattern is a real HTML page carrying the summary, the key findings and the metadata, with the PDF offered from it — the page is what gets found, and the file is what gets downloaded.
Structure · What to fix before submitting anything

Reducing the surface rather than pushing harder

Every hour spent on submission tooling before the structural work is done is an hour spent making a queue longer. The order below is deliberate: each step reduces the number of addresses competing for attention, so that the ones you care about are reached sooner.

StepActionTypical effect on a trilingual site
1Remove parameter and filter combinations from the crawlable setOften the largest single reduction
2Return a proper 404 for untranslated documents instead of falling backEliminates cross-language duplicates
3Make each language version canonicalise to itselfStops two versions asking to be dropped
4Check that language pairing is declared in both directionsUnreciprocated declarations are discarded
5Split one large sitemap into one per languageMakes coverage gaps visible per version
6Link new documents from an index page within one clickDiscovery without waiting for a submission

Step five earns its place for a reason that is not obvious. A single sitemap listing twenty thousand addresses tells you almost nothing when coverage is poor. Three sitemaps, one per language, immediately show whether the Dutch version is being taken up at the same rate as the French one — and in practice it usually is not, because the language with fewer inbound links is crawled less.

Step six is the cheapest and the most neglected. An internal link from a page that is already crawled regularly is the most natural discovery mechanism there is, and it costs nothing. A document reachable only from a sitemap entry is a document that depends entirely on a queue.

Publish in a pattern, not in bursts. Ten documents released on one afternoon compete with each other for the same crawl window. The same ten spread across a fortnight, each linked from an index page, are almost always taken up more completely.
Routine · A schedule that fits a publishing calendar

What to do weekly, monthly and once

Indexing work rewards regularity far more than intensity. The routine below assumes an organisation publishing a handful of documents a week in three languages, and takes under an hour a month once the one-off work is behind you.

Once

Reduce and declare

Work through the six structural steps, then submit the full sitemap set as a single import.

  • Use the bulk batch for this, not for routine work
  • Record the date — it becomes your baseline
Weekly · 10 minutes

New documents

Submit the week's publications, all language versions together, and link each from its index page.

  • Well inside the daily budget
  • Check the found counter a week later
Monthly · 20 minutes

The failure log

Read the failed addresses and sort them by directory before reading them individually.

  • Clusters mean templates
  • Scatter means infrastructure
Quarterly · 30 minutes

Coverage per language

Compare the found counts across the three sitemaps and look for a version falling behind.

  • A lagging version usually lacks internal links
  • Fix linking before resubmitting

Where an automated campaign is running, the indexing side stops being a separate chore: newly published addresses go into the queue as part of the workflow, and the assistant in the campaign area accepts URL lists in batches alongside keyword lists. That is a convenience rather than a different outcome — the same submissions, made without anyone remembering to make them. Both the campaign automation tiers include it, at 149 USD per month per domain for the automated level and 500 USD for the level that adds a team of specialists, developers and writers to the same tooling.

Expect the first measurable movement four to eight weeks after the structural work, not after the first submission. Indexing decisions are revisited on a schedule you do not control, and a page that was declined in March is not reconsidered in April merely because you asked again. What changes the answer is a changed page, or a site with less competing for the same attention. The indexing log and its per-URL history are how you tell which of the two is happening.

Frequently asked questions

How long should I wait before deciding a page will not be indexed?

Two to three weeks after publication and submission, provided the page is linked internally. Beyond that, waiting is not a strategy — check the log to see whether it was fetched at all, because a page never crawled and a page crawled and declined need opposite responses.

Does submitting the same URL repeatedly help?

No. A submission asks for the address to be looked at; if it has already been looked at and declined, asking again produces the same decision. Repeated submission of unchanged pages is also a poor use of a daily budget that has better claims on it.

Should our three language versions share one sitemap or have their own?

Their own, for diagnosis rather than for crawling. A search engine treats the entries identically either way, but separate files let you see coverage per language, and a version being taken up more slowly than the others is a finding you would otherwise never make.

We publish mainly PDFs. Is that a problem?

It is a handicap rather than a blocker. PDFs can be indexed but rank poorly, carry no internal links onward and give the reader nothing to navigate from. Publish an HTML page with the summary and findings, and attach the file to it.

Our archive has thousands of old documents. Should they all be indexed?

Probably not, and deciding so deliberately is one of the more useful things you can do. Documents that are superseded, procedurally obsolete or duplicated across languages consume crawl attention that your current publications need. Keep them reachable for people; remove them from the crawlable set.

Where the effort actually pays

What none of this delivers. No submission mechanism guarantees indexing, no tool can tell you why a specific page was declined, and no reporting layer knows more about the decision than Google does. What the tooling provides is submission at scale and an evidential record of what the crawlers did — useful, bounded, and frequently oversold elsewhere.

The organisations that solve this rarely do so by finding a better submission tool. They do it by noticing that they had built twenty thousand addresses where four thousand would have served, and by making the newest document easier to reach than the oldest archive filter. The tooling then does what tooling is good at: it makes the mechanical part fast and it keeps the receipts.

For a trilingual publisher the calculation is simply sharper. Three versions of everything means the cost of a structural mistake is paid three times, and the benefit of fixing one is collected three times as well. That is an unusually good return for work that mostly consists of deciding what does not need to be crawled. If you want the adjacent reading, our technical audit notes and the other articles here cover the measurement side.

To see where your own site stands, connect the property, submit the sitemap set and then leave the log alone for a fortnight before drawing conclusions: open the dashboard and start an indexing job. The useful number will not be how many addresses you submitted. It will be the difference between that figure and how many were actually found.