The Whale Creative

Web Design

How Does a Website Actually Rank in Search Engines? The Technical Side of SEO

· Author: The Whale Creative · 20 min read

In short

  • A website gets into search results in three stages: crawling, indexing and ranking; if either of the first two is blocked, content quality never even gets evaluated.
  • If a page doesn't appear, first check in Search Console's URL Inspection whether it's indexed at all; keyword work on a page that isn't indexed produces nothing.
  • Good Core Web Vitals means LCP under 2.5 seconds, INP under 200 milliseconds and CLS under 0.1, measured on real-user data at the 75th percentile.
  • Blocking a page in robots.txt while also adding noindex doesn't work: the bot can't fetch the page, so it never reads the noindex directive and the page can stay indexed.
  • Google removed FAQ rich results on May 7, 2026; FAQPage markup is still valid, but it no longer earns any visual benefit in search results.

Crawling, indexing and ranking step by step: Search Console diagnosis, Core Web Vitals thresholds, schema, mobile-first and a technical SEO audit.

How Does a Website Actually Rank in Search Engines? The Technical Side of SEO

A beautiful website that nobody finds is a portfolio piece, not a business tool. Search engines don't judge design taste — they judge structure, speed, and clarity. Here is what actually happens behind the scenes before a site ever reaches page one.

We've written this as a work list rather than a glossary. In each section you'll find where to look inside which tool, why a particular mistake keeps getting made, and when the advice in that section doesn't apply to you.

How Do Search Engines Actually Rank a Page?

A search engine puts your page through three stages: crawling, indexing, and ranking. If any one of them breaks, nobody ever gets to the question of whether your content is good — the page simply never entered the race.

Crawling is the bot being able to fetch your page at all. A robots.txt block, a server error, or content that only appears after JavaScript runs will all cause trouble here. Indexing is that fetched page being accepted into the search engine's database; duplicate content, a misapplied canonical tag, or a stray noindex directive close the door at this stage. Ranking only begins once the first two are clean.

Your diagnostic order should follow the same sequence. If a page isn't showing up, start with the URL Inspection tool in Google Search Console and check whether it's indexed at all; if the answer is no, keyword work is pointless. Skipping this order is where SEO budgets most often get wasted.

How do you run the diagnosis, step by step?

Say a service page has been live for three weeks and appears for nothing. Before panicking and rewriting the content, work through this order.

1. Paste the page's full URL into the URL Inspection box in Search Console and check whether it says the page is on Google. 2. If it isn't indexed, run the live test with "Test live URL"; this tells you whether Google can fetch the page at all. 3. If it can't, check for a robots.txt block, a server error, or a redirect chain. 4. If it can, open the rendered HTML and look for your main copy; if it isn't there, the problem is content generated by JavaScript. 5. If the page is indexed, move to queries: check the Performance report for the impressions that page receives. 6. If there are no impressions at all, the problem isn't technical, it's relevance — the page is about something nobody searches for.

Checklist: baseline visibility

  • Does Search Console report the page as indexed?
  • Does the page's canonical tag point at itself?
  • Does at least one internal link point to the page?
  • Is the page included in sitemap.xml?
  • With JavaScript disabled, is the main copy still visible?

Why Doesn't a Page Get Indexed?

Most of the time a page stays out of the index not because of some mysterious penalty but for one of four concrete reasons: it can't be reached, it's a duplicate, it was judged not worth including, or you excluded it without realizing. The Pages report in Search Console already tells you which one applies, by name.

Here's what the common reasons mean in practice. "Discovered — currently not indexed" is usually a discovery or value problem; add internal links and strengthen the content. "Crawled — currently not indexed" means Google saw the page and didn't consider it worth indexing; that's a quality judgment, not an error. "Alternate page with proper canonical tag" isn't a problem at all — it shows duplication is being handled correctly. "Duplicate, Google chose different canonical than user" tells you your canonical preference was overruled, and that one deserves investigation.

A common mistake: combining robots.txt with noindex

The most frequent technical mistake is a team that wants a page kept out of search blocking it in robots.txt *and* adding noindex to it. That doesn't work, because a bot that can't fetch the page can never read the noindex directive on it; the page can stay in the index unread. The fix is simple: remove the robots.txt block, leave noindex in place, and wait for the page to be crawled so the directive is seen. Blocking crawling doesn't guarantee de-indexing; more often it produces the exact opposite.

When do you want a page left out of the index?

Not every page belongs in the index. Internal search result pages, filtered listing URLs, thank-you pages, print versions, and account pages containing personal data are deliberately excluded. Keeping those URLs indexed buys no visibility; it generates hundreds of near-identical thin pages that dilute the ones that matter.

Speed Isn't a Detail, It's a Ranking Factor

Google measures your site through Core Web Vitals — how fast content appears, how quickly it becomes interactive, how much it visually shifts around while loading. A slow site doesn't just frustrate visitors; it gets actively pushed down in rankings. This is why we build on modern, lightweight foundations and ship compressed, optimized images by default, instead of treating performance as a launch-day afterthought.

Core Web Vitals is currently made up of three metrics, each with a clear "good" threshold:

  • LCP (Largest Contentful Paint): how long the main content takes to appear. Good is under 2.5 seconds.
  • INP (Interaction to Next Paint): how quickly the page responds to an interaction. Good is under 200 milliseconds. INP replaced the older FID metric on March 12, 2024.
  • CLS (Cumulative Layout Shift): how much the page jumps around while loading. Good is under 0.1.

One detail matters more than people expect: these numbers come from real user data, not a lab test, and they're assessed at the 75th percentile. Three quarters of your visitors have to clear the threshold. Your own fast laptop loading the site smoothly proves nothing.

The three most common sources of a slow page

1. Oversized, unscaled images — fixed with modern formats, correct dimensions, and lazy loading. 2. Too many third-party scripts — every chat widget, heatmap, and pixel competes for the main thread. 3. Images and ad slots without reserved dimensions — content shifts as they load, which wrecks CLS.

How do you fix a slow page, step by step?

Picture a boutique hotel's homepage: a large hero image at the top, a booking widget directly beneath it, an embedded map inside the page, and four separate tracking scripts. On mobile it's visibly slow. Work through it in order.

1. Open the page in PageSpeed Insights and read the real-user data at the top first; the lab score comes second. 2. Identify which single metric fails its threshold and work only on that one; chasing all three at once wastes time. 3. If LCP is the problem, check the hero image's actual dimensions, convert it to a modern format, and remove lazy loading from that one image. 4. If CLS is the problem, give images and embedded frames explicit width and height, and make sure font swaps don't push text around. 5. If INP is the problem, list every third-party script, ask what business result each one produces, and remove the ones that can't answer. 6. Wait four weeks after the change; real-user data is reported over a 28-day window, so results don't appear immediately.

A common mistake: chasing the score

Plenty of teams treat the out-of-100 PageSpeed Insights number as the goal and strip functionality out of a page to raise it. That's the wrong target, because what counts for ranking is field data collected from real users, not the lab score; a page scoring 100 can still fail in field data. Set the goal as "all three metrics inside their thresholds" rather than "a higher score," and confirm every improvement against real-user data.

Search Engines Read Structure, Not Just Words

A page full of great copy wrapped in generic, meaningless tags tells a search engine almost nothing about itself. Semantic HTML (proper heading hierarchy, article regions, landmark structure) and structured data (schema markup) tell it exactly what it's looking at — a service, a blog post, an FAQ, a product. This same layer also determines whether AI systems can parse your page cleanly at all.

The current standard for structured data is the Schema.org vocabulary implemented as JSON-LD. For an agency or local business site, the useful types are few, and you don't need more than these:

  • Organization or LocalBusiness — who you are, where you are, how to reach you.
  • Article or BlogPosting — the piece, its author, its publication date.
  • Service or Product — what you actually offer.
  • BreadcrumbList — where the page sits within the site.

Set your expectations correctly, though: Google removed FAQ rich results on May 7, 2026, followed by the FAQ appearance report and Rich Results Test support in June 2026 and Search Console API support in August 2026. FAQPage markup didn't become invalid and leaving it in place causes no harm, but it no longer buys you anything visual in the results. The feature had in any case been limited to government and health sites since August 2023, so for most businesses nothing practical was lost.

How do you build a heading hierarchy?

Headings aren't a visual choice; they're the page's table of contents. A page carries a single h1 that states its subject. Main sections become h2, their subdivisions h3; levels don't get skipped, and heading tags aren't used to make a line of text look bigger. A correctly built hierarchy is what lets both screen readers and answer engines break the page into sections.

Don't do this: a "Services" heading followed by six different services listed at the same level as unlabeled paragraphs. Do this: make "Services" an h2, give each service its own h3, and make the first sentence under each h3 a one-sentence definition of that service.

A common mistake: shipping unvalidated schema

Whether it's added by hand or by a plugin, schema markup frequently ends up describing things the page doesn't contain — a rating that doesn't exist, a wrong date, a price that never appears on screen. That isn't an advantage, it's a risk; marking up information not visible on the page breaks structured data policy. Check it with the Rich Results Test and a schema validator before it ships, and mark up only what genuinely appears on the page. Schema that describes something the page doesn't contain is worse than no schema at all.

Meta Tags and Titles: Your Business Card in the Results

The title and the meta description are the only two things most people ever see before they decide whether to click. A generic, copy-pasted title gets ignored; a specific one earns the click. And a page that collects impressions without clicks is telling you something on its own: the promise in the title isn't matching what the searcher came for.

The pattern that works for a title tag is simple: what it is, who it's for, the brand name. Judge the length by whether it gets truncated on screen rather than by a character count; Google truncates by pixel width, and will rewrite a title it considers unhelpful using other text from the page.

Do this:

  • Give every page one unique title and one unique description.
  • Put the page's actual subject in the first few words of the title.
  • Write the description as a summary of what's on the page, not as a promise.

Don't do this:

  • Don't repeat the same title across the whole site.
  • Don't string keywords together separated by commas.
  • Don't promise something in the description that isn't on the page — that visitor bounces straight back.

What does this look like at sentence level?

Don't write: "Home | Corporate Web Design, SEO, Social Media, Digital Marketing, Advertising". Write: "Hotel Website Design and Booking Engine Integration — The Whale Creative". Don't write a description like: "Get in touch for solutions that take your brand to the next level." Write: "We combine design, multilingual content, and booking engine integration for hotel websites into one project, delivered in four stages."

Checklist: titles and descriptions

  • Is the title different from every other page on the site?
  • Does the page's subject appear in the first five words of the title?
  • Does the description summarize something genuinely on the page?
  • In the Search Console Performance report, is this page getting impressions but no clicks?
  • Is Google rewriting the title itself? If so, look at which on-page text it prefers.

Mobile-First Isn't Optional

Google indexes and ranks your mobile version first, not your desktop one. As of July 5, 2024 it completed the migration by switching crawling of all sites to the smartphone Googlebot, which means the phone version of your site is now the version that counts. A site that merely "works fine" on mobile, rather than being genuinely designed for it, already starts from behind.

The practical consequence is harsher than most brands expect. Content you don't show on mobile is content that Google doesn't see at all. A section that exists on desktop but is hidden on small screens, structured data that only loads on the desktop build, internal links that appear only on wide viewports — all of it is missing from the version that gets indexed.

The checklist is short:

1. Is the same content present in both versions? 2. Do titles, meta tags, and structured data load on mobile too? 3. Are tap targets comfortable to hit, and is body text readable without zooming? 4. Does a cookie banner or pop-up cover the content? 5. Are internal links tucked inside a mobile menu present in the page source?

A common mistake: "simplifying" content on mobile

Design teams often treat hiding sections as the natural way to lighten a mobile experience. The usual result: a service description that runs three paragraphs on desktop becomes one sentence on mobile, and that one sentence is what gets indexed. Collapse rather than delete — put the text in an expandable section so it stays in the source without taking up screen space.

Sitemaps, robots.txt, and Being Discoverable

None of the above matters if search engines can't find your pages in the first place. A clean sitemap.xml and a correctly configured robots.txt are the unglamorous plumbing that make sure every important page actually gets crawled and indexed — and that pages you don't want indexed don't leak into search results by accident.

One point gets confused constantly. Robots.txt controls crawling, not indexing, and treating the two as the same thing produces exactly the outcome you were trying to avoid. If you need a page to stay out of search results for certain, the right tool is a noindex directive — but if you also block the page in robots.txt, the bot can never read that directive.

A correct setup looks like this:

  • sitemap.xml lists only canonical URLs that return a 200 status and that you want indexed.
  • The sitemap location is declared in robots.txt and submitted through Google Search Console and Bing Webmaster Tools.
  • Redirected, deleted, and low-value URLs such as tag archives are kept out of the sitemap.
  • Every page carries a self-referencing canonical tag.
  • Multilingual sites declare hreflang relationships between language versions.

What's the order when launching a new site?

1. Before launch, confirm the blanket robots.txt block from the staging environment has been removed; this is the single most common launch accident. 2. Search for leftover noindex tags on individual pages and clear them. 3. Generate sitemap.xml, then actually open it and read the URLs inside. 4. Verify ownership in Search Console and Bing Webmaster Tools and submit the sitemap. 5. Inspect your five most important pages individually with URL Inspection and request indexing. 6. Watch the Pages report for the first two weeks and read the reason given for every excluded URL.

Content and Search Intent: Where Technical Work Stops Being Enough

Technical groundwork gets your page considered; it doesn't get it ranked. Ranking comes down to whether the page serves the searcher's intent better than the alternatives — and intent matters more than the keyword itself.

The same topic can carry four different intents: informational ("what is a brand identity"), comparative ("agency or freelancer"), local ("web design Antalya"), and transactional ("get a web design quote"). Compressing all four into one page means being mediocre at all of them.

Three questions to ask before writing a page:

1. Is the person searching this after an answer, a comparison, or a service? 2. What format do the pages currently ranking take — a guide, a list, a service page? 3. What concrete information does our page have that theirs doesn't?

If you can't answer the third one, publishing that page won't earn a ranking. In most cases, improving a page you already have beats adding another thin one.

How do you verify intent?

Look at the results page instead of guessing. Search your target query in a private window and note the format of the top ten results: how many are guides, service pages, lists, videos. That distribution tells you which format Google considers appropriate for the query. If you want to rank a service page for a query where nine of the top ten are guides, your problem isn't the quality of your page, it's its format.

A common mistake: writing several pages for one intent

Teams routinely create three separate pages for "web design," "web design services," and "professional web design." Because all three carry the same intent, the pages compete to substitute for one another and none of them ranks consistently. Merge them into the strongest one, redirect the rest to it, and handle the variations as sections rather than as separate titles.

How Do You Run a Technical SEO Audit?

A technical SEO audit is a structured review that checks your site's crawlability, indexing status, and performance in that order. You can get surprisingly far with Google's own free tools before paying for anything.

Work through them in sequence:

1. Google Search Console, Pages report: how many pages are indexed, and for what reason the rest aren't. 2. URL Inspection: check key pages individually and read the rendered HTML Google actually sees. 3. Core Web Vitals report and PageSpeed Insights: read the real-user data alongside the lab data. 4. Lighthouse: a fast per-page scan for performance, accessibility, and best practices. 5. Performance report: queries with impressions but no clicks show you exactly which titles and descriptions to revisit. 6. A manual check: view a page with JavaScript disabled — is the main copy still there?

Then sort what you find by impact. A service page that isn't indexed always matters more than a few hundred milliseconds of speed.

How do you prioritize the findings?

An audit usually leaves you holding a list of thirty or forty items, and doing all of them at once isn't possible. Judge each item with two questions: does this affect a page that generates revenue, and does fixing it change visibility directly? Items that answer yes to both go into the first week, items answering yes to one go into the month, and the rest go into the next development cycle. An audit report that hasn't been sorted by impact is an audit report nobody implements.

When's Technical SEO Not Your Priority?

It's worth saying that technical SEO isn't equally urgent at every business. This work pays off when search-driven demand already exists and your site has the content to meet it.

In the situations below, spending the budget elsewhere is the better decision.

  • The site is already technically clean and the problem is the content. Pages that are indexed but get no impressions have a relevance problem, not a speed problem.
  • Your demand comes entirely through referrals, tenders, or an existing client network. Winning rankings in a space with no search volume produces no revenue.
  • The site is being rebuilt in the coming months. Optimizing infrastructure you're about to remove is wasted effort; the job here is writing the technical requirements for the new build.
  • The product or service isn't defined yet. Until you know what you sell, you can't decide which query you should match.
  • You have no measurement in place at all. Without Search Console and analytics, the effect of every improvement is unmeasurable, so setup comes first.

How We Approach It

Every project we build is engineered with this layer in mind from day one, not patched in afterward. Aesthetics without infrastructure simply doesn't get seen.

In practice that means we discuss heading hierarchy and page structure during design, because a heading isn't just large type. Images are resized and converted to modern formats before launch, not after. Every project's delivery list includes sitemap.xml, robots.txt, canonical tags, hreflang for multilingual sites, and baseline structured data. Once a site goes live we check its indexing status in Search Console — because "published" and "findable" aren't the same thing.

We keep one habit after handover too: for the first two weeks post-launch we watch the Pages report, read the reason given for each excluded URL, and fix what needs fixing. We don't make performance promises we can't verify either; we measure what is measurable and let the rest show over time.

We don't consider "we delivered an SEO-friendly site" a finished sentence; the job isn't done until the pages have entered the index.

Frequently Asked Questions

How long does SEO take to show results?

Giving a fixed timeline wouldn't be honest; it depends on the site's age, the competition, and the kind of problem being fixed. Broadly, indexing problems show improvement quickly, while work that requires content and authority is measured in months.

If I fail Core Web Vitals, will my rankings drop?

Core Web Vitals are among the ranking signals but they aren't decisive on their own. Relevance and content quality carry more weight; speed is the layer that separates closely matched competitors and directly affects conversion.

Is a one-page site a disadvantage for SEO?

Usually yes, because you can't offer a separate URL and a separate title for each search intent. If your services are searched with different queries, giving each service its own page and URL is a clear advantage.

Is blogging really necessary for rankings?

What's necessary isn't a blog; it's answers to questions that haven't been answered well. A small number of deep pages that genuinely resolve what your customers ask outperforms a steady stream of shallow posts.

Does submitting a sitemap guarantee my page gets indexed?

No. A sitemap is a discovery aid, not an instruction; the search engine finds the page, then decides on its own quality assessment whether to index it. When pages stay out of the index, the cause is usually thin originality or duplicate URLs.

Can I block a page in robots.txt and add noindex at the same time?

That combination doesn't work, and it's a common mistake. A bot that can't fetch the page can't read the noindex directive on it; for a page to leave the index it has to be crawlable and the directive has to be seen. Remove the block first, leave noindex in place, then wait for the crawl.

Will a JavaScript-driven site have SEO problems?

Not necessarily, but your risk is higher because content only exists after rendering. The practical check: load the page with JavaScript disabled, or read the rendered HTML in Search Console. If the main copy, headings, and internal links aren't there, you need server-side rendering or pre-rendering.

Do AI search engines require a separate technical setup?

No separate setup is needed; generative search features run on the same crawling and indexing infrastructure. The only differences are which AI crawlers you permit in robots.txt and whether your content is written in a quotable form.