ParitLAB
← Lab Notes

Guides · 2026-09-15

Behind a bilingual website: keeping readers and Google on the right page

Published by ParitLAB
Field notes from designing, building and testing our products

เว็บไซต์สองภาษา, การค้นหา, ประสบการณ์ผู้ใช้

A bilingual site contains intentionally similar information, but inconsistent URLs and metadata can leave a search engine unsure which page is Thai, which is English, and which URL should represent the content. The problem becomes more visible on a PHP site that uses a query string such as ?page=notes&lang=en rather than language folders.

ParitLAB uses one controller for the requested page and language. We centralized URL generation and made the page title, canonical link, language switcher, sitemap and structured data follow the same rule. That reduces cases where internal navigation points to one URL while the canonical tag declares another.

Define one canonical form

Every public page has an absolute HTTPS canonical on the non-www domain. Parameters that identify different content, such as an article slug or software slug, remain in the canonical. Search, category and internal-state parameters do not create a new canonical page unless they represent independently useful content.

Requests through HTTP, www, or /index.php are redirected to the same format. Consolidating these variants keeps internal links, analytics and indexing signals from being divided among equivalent addresses.

Connect language equivalents with hreflang

Each public page declares alternates for th, en and x-default. Thai and English pages refer back to one another. On a detail page, the article or product parameter must also remain in the language-switch link. Otherwise, selecting English from an article returns the reader to the article index, creating a broken reading flow and an inaccurate language relationship.

Let the sitemap reflect published data

Our sitemap is generated from the core public pages, published projects and published articles. Every entry has both language versions and alternate links. Drafts, account pages, administrator pages and authenticated views are excluded.

One practical detail is that robots.txt must advertise the URL that actually serves the sitemap. A query URL can be used directly; a physical /sitemap.xml file is not required as long as the endpoint is accessible, returns XML and uses the appropriate content type.

Separate public content from application screens

Content pages such as Home, Software, Store, Lab Notes, About, Terms and Privacy use index,follow. Login, password recovery, the client portal and administrator pages use noindex,nofollow and are excluded through robots.txt. Filtered Lab Notes results also use noindex when a query, category or tag is present, avoiding a large set of low-value combinations.

Keep structured data consistent with visible content

Articles use BlogPosting data for the headline, description, language, publication date, author and publisher. Each value comes from the article shown on the page. We do not add review scores or facts that a reader cannot verify. Structured data should clarify a page, not make a separate claim about it.

Pre-release checklist

Technical SEO cannot replace useful content, but it helps original work get discovered and grouped correctly. Centralizing the URL rule also makes new pages safer: an editor can publish an article without remembering to update several unrelated tags by hand.