UNmiss Blog

TED's Real Archive Isn't the Part You Watch

TED publishes 117,053 URLs and its curator-approved talks are 6.5% of them. We counted all 29 sitemap files and found the two sets never overlap.

Say "TED" and a picture arrives before you finish the word: a red circle, a dark stage, a speaker walking to a mark. That picture is a video library, and it is what the brand is known for.

So we did to ted.com what we have done to every site in this series: opened the sitemap index, fetched every child file, and counted what TED actually asks search engines to crawl.

The library is in there. It is just not what the sitemap is mostly made of. What TED submits is overwhelmingly an archive — year-bucketed files of talk pages no curator ever put on the front of the site. The famous part is a small slice, and the rest is a long tail the brand's name quietly holds up.

Results at a glance

117,053
URLs in the sitemap
12.7:1
Archive pages per curated video
7.8
Referring domains per page
The short version

TED's public video library is 6.5% of its sitemap. The other 97,014 talk URLs live in year-bucketed files — a separate, far larger archive that shares no URLs at all with the curated set, because TED's generator lifts the curated pages out of the year buckets by design.

The thing that makes this work is not the architecture. It is that TED earns 7.8 referring domains per published page, the third-highest figure in this entire series, so even the long tail sits under a very large canopy of links.

What the sitemap actually contains

The root index points at 29 child sitemaps. Fetched end to end, they hold 124,676 raw URL entries. That is not the number of pages, and a crawler reporting it would be wrong twice over.

The first correction is a deliberate duplicate. TED submits the same curated URL set in two files: one plain, one wrapped in video schema markup with thumbnails, titles and descriptions. The sets are identical, URL for URL, not merely equal in count. That is a video-search play, not a mistake, and those pages count once.

The second is a malformed file. One child is declared as a page list, but its entries point at other sitemap files that already appear in the root index. Read literally or sensibly, TED publishes 117,053 or 117,051 URLs. Our two independent counts landed either side of that gap, and nothing else separated them.

A talk page on ted.com showing the video player, speaker name and transcript navigation
What to notice: curated and year-bucketed archive talks share the same /talks/ URL pattern. Nothing in the address tells you which side of the split a page came from — only the sitemap does.

The label that did not survive

Our first pass reached for the obvious story. Call the curated file "official TED talks", call the year buckets "TEDx", and you have a neat dichotomy: the famous stage on one side, the licensed events on the other.

The independent recount killed it. Someone pulled a 40-page sample out of the curated set and read the pages. A fifth were TEDx talks. Another fifth were TED-Ed animated lessons. The curated file is not an official-versus-licensed boundary — it is an editorial one. TED's curators reach into TEDx and TED-Ed and pull the best of both onto the front of the site. The split is not organisational; it is quality control, and that changes what you should copy.

It also deflates a claim we liked. The two sets share exactly zero URLs, which sounds remarkable until you add them up and find they cover TED's entire talk namespace. They partition it by construction: the generator moves a page out of its year bucket when a curator promotes it. Zero overlap was never a coincidence.

Why every page carries links

Here is why an archive this size does not drag the site down. TED earns 7.8 referring domains for every page it publishes. In this series that is beaten only by calculator.net at 143 across 222 pages and Healthline at 14.4 across 42,567 — both far smaller sites. TED clears MDN's 3.9, and it is more than 8 times GOV.UK's 0.906. The large publishers we have measured fall off a cliff at this scale: itch.io sits at 0.318, Eventbrite at 0.148, Meetup at 0.069, Rotten Tomatoes at 0.063.

Our tool puts ted.com's domain rank at 96, level with PyPI, Meetup and Healthline, behind MDN's 100 and Eventbrite's 98. The referring-domain figure is the more interesting one. Rank tells you how strong a domain is; domains per page tells you whether that strength reaches the pages nobody talks about.

UNmiss Backlink Analyzer report for ted.com showing domain rank, referring domains and the dofollow to nofollow split
What to notice: 913,482 referring domains, and 95.9% of the links marked dofollow. A nofollow share that small usually means editorial citations rather than profiles and comment fields.

The year that is missing

The year buckets run from 2006 to 2025, and the run is not contiguous. There is no file for 2008. Every other year has one, including years holding a single URL.

Small as it is, that gap is the most useful thing in the crawl, because it is the kind of defect that only shows up when you enumerate. The index still validates. Every file still returns a clean response. Nothing is broken enough to alert on. The pages simply are not submitted, and the only way to notice is to lay the filenames out in order.

The same crawl turned up 9 talk slugs carrying internal working markers — the sort of thing typed into a title field without expecting it to reach a public URL. All 9 return not-found now, but they were still listed for crawlers to fetch. A sitemap is a publishing surface, and it inherits whatever the content system hands it.

Built well, not built perfect

Running the site through our audit gives 83, with no critical issues, 3 warnings and 8 notices against 110 passing checks. That places TED level with Meetup and levels.fyi, under PyPI's 91 and MDN's 88, above Rotten Tomatoes at 80 and Healthline at 74.

Read as a report card that is a middling grade. Read as a fact about how the site is built, it says something more specific: nothing structural is wrong. Technical checks come back at 64 out of 67 and mobile is clean. The points come off speed, at 23 out of 29 — exactly where a video-first site pays.

One trap if you go looking yourself. Most of TED's child sitemaps carry a compressed file extension, but the server hands them over already decompressed to an ordinary browser. Code that unconditionally decompresses anything ending in .gz falls over here. Check the leading bytes, not the name.

Copy this in an afternoon

None of this requires TED's brand. It requires deciding what your sitemap is for.

1. Split curated from archive in the file structure, not just in the navigation. TED keeps its promoted pages in their own sitemap and its long tail in date-bucketed files, so indexation on the pages you care about can be watched separately from the pages that exist for completeness.

2. Submit your video pages twice, on purpose. TED lists the same curated URLs a second time wrapped in video schema markup. That is not duplication in any sense that matters — it is one page set offered to a second search surface with the metadata it wants.

3. Lay your sitemap filenames out in order and look for gaps. A missing year, a skipped shard, a bucket that never generated: none of these produce an error anywhere. They are visible only as an absence, and only if you list the files.

4. Audit the slugs your CMS is publishing, not just the pages. Internal markers and draft labels end up in URLs, and URLs end up in sitemaps. Grep your own sitemap for the words your team uses when a page is not meant to ship.

What you cannot copy

You cannot copy the canopy. TED's 913,482 referring domains were not earned by the 97,014 archive pages — they were earned by a handful of famous talks and the brand attached to them, and they now sit over everything on the domain. That is what makes a 12.7:1 archive-to-curated ratio safe to run. Without that canopy, the same ratio is just thin pages under no protection. Get the link equity first, then let the archive grow underneath it.

And here is what we could not check:

Our two counts differ by 2 URLs. One file is declared as a page list but contains pointers to other sitemap files. Read literally it yields 117,053; read sensibly, 117,051. We publish the larger figure and flag it rather than quietly picking a side. This is a full parse of all 29 children, not an estimate.

The composition labels rest on samples, not on a full classification. We read 40 pages from the curated set and 28 from the year buckets, plus 9 detail pages drawn from spread positions across the files. That is enough to say the year buckets are dominated by TEDx events and the curated set is mixed. It is not enough to give a percentage for each type across all 117,053 URLs, and we do not.

The year curve is not a publishing trend and we will not present it as one. The buckets swing from a single URL to more than 21,000, and the most recent two hold a different kind of content entirely — TED-organisation and podcast uploads rather than TEDx talks. Plotting those years against each other would compare unlike things and produce a dramatic, meaningless chart.

Our backlink tool's headline total does not perfectly reconcile with its own dofollow and nofollow split. The two figures are within a few per cent of each other, inside the tolerance we accept, but they are not the same arithmetic. Treat the total as an order of magnitude and the referring-domain count as the number that carries the argument.

A sitemap is a request, not an outcome. Everything here describes what TED asks search engines to crawl. We did not measure what is indexed, what ranks, or what any of it earns in traffic. Liveness is good — 60 randomly sampled archive URLs all returned successfully, so this is no stale file submitted out of habit — but crawlable is not valuable.

Free, no account
See what your sitemap is really submitting

TED's sitemap hides a missing year, a deliberate duplicate and slugs that were never meant to ship. Yours almost certainly hides something. Find out before a search engine does.

  • Checks sitemap structure, indexability and crawl signals
  • Separates critical faults from warnings and notices
  • Scores on-page, technical, speed and mobile independently
Run a free website audit

Frequently asked questions

How many pages does ted.com publish in its sitemap?

Between 117,051 and 117,053, depending on how you treat one malformed file. The raw entry count across all 29 children is 124,676; the gap is a deliberately duplicated set of 7,623 curated URLs plus 2 entries pointing at sitemap files, not pages.

Why do your two counts disagree?

One child file is declared as a list of pages, but its 2 entries are addresses of other sitemap files that already appear in the root index. Our first count excluded them as not being pages; our verification count included them literally. That is the entire discrepancy: 0.002% of the total.

Is the curated video library really only 6.5% of the site?

Yes. 7,623 curated URLs out of 117,053 is 6.51%. The 97,014 year-bucketed talk URLs work out at 12.7 archive pages for every curated one.

Are the year-bucketed talks all TEDx?

Overwhelmingly, but not entirely. All 28 pages we sampled from the 2006 to 2023 buckets were TEDx events. The 2024 and 2025 buckets differ: they hold 56 and 70 URLs and carry TED-organisation and podcast content, together 0.13% of the bucket.

Is the curated set the same as official TED stage talks?

No, and this is where our first framing was wrong. In a 40-page sample of the curated set, 8 pages were TEDx talks and 8 were TED-Ed animated lessons. The curated file is an editorial selection from across TED's output, not a list of main-stage talks.

Why is a year missing from the archive?

We do not know. The year files run from 2006 to 2025 with no file for 2008, while years holding a single URL do have one. It is an absence rather than an error, which is why it is easy to miss.

How strong is ted.com's link profile?

Our tool reports a domain rank of 96 and 913,482 referring domains, with 100,896,462 links marked dofollow against 4,331,141 nofollow — a 95.9% dofollow share. The reported total of 101,458,916 does not reconcile exactly with that split: a limitation of our tool, not a finding about TED.

Does TED submit the same pages twice?

Yes, deliberately. The 7,623 curated URLs appear in a plain sitemap and again in a video sitemap carrying full schema markup — thumbnails, titles, descriptions. We verified the two URL sets are identical, so we counted them once.

The lesson TED offers is not about scale. Plenty of sites in this series publish more pages, some by a factor of 40, and most earn a fraction as many links per page.

It is about knowing which of your pages are load-bearing. TED draws a hard line between the pages a human chose and the pages that simply exist, and it draws it in the file structure where a search engine can see it. The curated 6.5% carries the brand. The other 93.5% is allowed to be an archive, because it is not pretending otherwise.

Most sites blur that line, submit everything in one pile, then wonder why their best pages are treated like their worst. If you have never looked at your own sitemap as a shape rather than a list, the free website audit is a reasonable place to start.

Measured on 29 August 2026. Page counts come from a full parse of ted.com's sitemap index and all 29 child sitemaps — no sampling, no estimation — cross-checked by an independent second count that reached 117,053 against the first count's 117,051.

Backlink figures and the audit score come from the UNmiss Backlink Analyzer and UNmiss Website Audit, run the same day with the settings used for every other site in this series. Page-type labels come from child sitemap filenames confirmed against sampled pages.

Blog