
Technical SEO: An Audit You Can Run Yourself
Keywords and good content only work if search engines can crawl your pages, index them, and render them the way visitors see them. A technical audit checks that plumbing. This guide walks through one you can run yourself in an afternoon with a browser, Search Console and a terminal, then turn into a short list of fixes ordered by what actually matters.
If on-page basics such as titles, headings and internal linking are still new, read the SEO optimization tutorial first. This post assumes those are in place and looks at the layer underneath.
Technical SEO answers three questions in order. Can a crawler reach the page? Is it eligible to be indexed, and is it the version you want indexed? Does it load and behave well enough for people to use? The order matters, because the questions stack: a blocked page cannot be indexed no matter how fast it is. Be honest about the ceiling too. Technical work removes obstacles, it does not create demand for content nobody wants.
Step 1: Check crawlability
Open https://yoursite.com/robots.txt in a browser. If nothing loads, that is fine: an absent robots.txt means everything is allowed. If something loads, read every line.
User-agent: *
Allow: /
Disallow: /cart/
Disallow: /search?
Sitemap: https://example.com/sitemap.xml
The most common failure here is a staging rule that shipped to production: a stray Disallow: / blocks the whole site. A one-line mistake that can take months to notice.
An important distinction that trips up a lot of people: Disallow means "do not crawl", not "do not index". A blocked URL can still appear in results, with no description, if other pages link to it. To keep a page out of the index you need a noindex meta tag or header, and the page must stay crawlable so that directive can be read. Blocking it in robots.txt prevents the crawler from ever seeing the noindex.
Then check the sitemap. It should list only canonical URLs that return status 200, with no redirects, no 404s and no noindex pages. Confirm it is referenced in robots.txt and submitted in Search Console.
Step 2: Confirm what is actually indexed
Search Console is the only source that tells you what Google itself sees. In the Pages report, look at the ratio between indexed and not indexed, then read the reasons given for the excluded group. Common ones:
- Crawled, currently not indexed. Seen and judged not worth storing. Usually thin or near-duplicate content, not a technical fault.
- Discovered, currently not indexed. Known but not yet fetched. Often weak internal linking to those pages.
- Alternate page with proper canonical tag. Working as intended, assuming the canonical points where you want.
- Duplicate, Google chose a different canonical. Your canonical was overruled, which means the signals disagree. Worth investigating.
- Excluded by noindex. Correct for tag archives and thank-you pages, alarming for anything else.
Then run the URL Inspection tool on a few pages: your homepage, a key landing page, a recent post. It reports the canonical Google selected, the last crawl date, and whether the page is eligible to appear. "Test live URL" fetches the page fresh, which is how you check a fix without waiting for a recrawl.
Step 3: Find duplicate URLs and check canonicals
The same content served at several addresses splits signals between them. The usual sources are mechanical: http and https, www and bare domain, trailing slash and no trailing slash, uppercase paths, /index.html alongside /, tracking parameters such as utm_source, and sort or filter parameters on listing pages.
Test those by hand. Every variant should redirect to one chosen form with a 301:
curl -sIL https://example.com/About/ | grep -i '^HTTP/\|^location:'
Then confirm each page declares its own address as canonical:
<link rel="canonical" href="https://example.com/blog/technical-seo-audit" />
Rules that keep canonicals out of trouble: use absolute URLs, point to a page that returns 200 rather than a redirect, keep it consistent with your sitemap and internal links, and never let a template emit the homepage as the canonical for every page. That last one happens more often than you would think, and it removes a whole site from results.
Parameters are the messier half. Where possible, have filtered and sorted views canonicalise to the clean base URL. Pagination is different: page two should be canonical to itself, not to page one, or you hide everything past the first page.
Step 4: Check how the page renders on mobile
Google indexes the mobile version of your pages. If your mobile template hides content the desktop one shows, the hidden content is effectively not there for ranking.
In your browser's developer tools, switch to a viewport around 390 pixels wide and check four things. Is there horizontal scrolling, which means something is overflowing? Are tap targets big enough to hit without zooming? Is body text readable without pinching? And is the same main content present, not stripped for a "mobile-friendly" layout?
Rendering deserves separate attention if your site is JavaScript-heavy. Use the URL Inspection live test and read the rendered HTML it returns. Content that appears in the browser but not in that output is content search engines may not have. Server-side rendering removes the question; with client-side rendering, verify rather than assume.
Step 5: Read Core Web Vitals in plain words
Core Web Vitals are three measurements of how a page feels to use:
- LCP, Largest Contentful Paint. How long until the biggest thing on screen, usually the hero image or headline, appears. Google's published threshold for "good" is 2.5 seconds or less.
- INP, Interaction to Next Paint. After a tap or click, how long before the page visibly responds. Good is 200 milliseconds or less.
- CLS, Cumulative Layout Shift. How much the page jumps around while loading. Good is 0.1 or less. Users describe this one as "I tapped the wrong button".
Two sources report these and they differ. PageSpeed Insights and the Search Console report show field data from real visits, which is what counts, but it needs enough traffic and covers a rolling window. Lighthouse gives a lab score instantly on your own hardware: useful for testing a change, unreliable as a verdict.
The causes are predictable. Slow LCP is nearly always an oversized hero image, or one loaded late through CSS or JavaScript instead of a plain image tag the browser can discover early. Poor INP is heavy JavaScript blocking the main thread. Layout shift is images without reserved width and height, or a web font swapping in and reflowing the text.
Step 6: Find broken links and redirect chains
Internal links pointing at 404s waste crawl attention and frustrate readers. Search Console's Pages report lists the 404s it has met; a crawler run against your own domain finds the rest, including ones nobody has clicked yet.
Redirects need two checks. First, chains: a URL that redirects to a URL that redirects again. Each hop costs time, and long chains sometimes stop being followed. The curl command from step three prints every hop. Second, targets: after a migration, a lazy rule often sends every old URL to the homepage, which is barely better than a 404 for the visitor and is often treated as a soft 404.
While you are checking status codes, verify that a genuinely missing page returns 404 rather than 200 with a "not found" message in the body. Pages that say "not found" while returning success are invisible to every tool you own.
Step 7: Validate structured data
Structured data describes your page in a format search engines parse directly, and it is what makes rich results such as breadcrumbs and review stars possible. JSON-LD in a script tag is the recommended form.
Start small: Organization or Person on the homepage, Article on posts, BreadcrumbList on anything nested, Product if you sell things. Test with Google's Rich Results Test for eligibility and the Schema Markup Validator for correctness. They answer different questions, so run both.
The rule that keeps you out of trouble: markup must describe what is visible on the page. Marking up ratings that appear nowhere is a manual-action risk rather than a shortcut. Rich results are never guaranteed either. Valid markup makes a page eligible; whether it shows is Google's call.
Turning findings into a fix list
An audit produces a long list. Sort it by blast radius rather than by effort. Site-wide problems that block crawling or indexing come first: a stray disallow, a broken canonical template, a sitemap that returns errors. Then issues affecting a whole set of pages, such as duplicate parameter URLs on your listings. Single-page issues and cosmetic warnings come last.
Fix in small batches and record the date. Changes take time to show up, because Google recrawls on its own schedule and field data updates over a rolling window. Weeks, not hours. Then re-run the same checks and compare.
Where to go next
- SEO Optimization for the on-page fundamentals this audit sits on top of.
- SEO tutorials for more search walkthroughs.
- UX and UI design for the layout and performance decisions behind Core Web Vitals.
Stuck on a step, or seeing a Search Console message you cannot place? Write to the desk with the URL and the exact wording of the report.
Comments
No comments yet. Be the first to share your thoughts.


