The wild west of polyglot docs sites#
2026 Oct 08
Under-the-hood, pigweed.dev is 3 separate docs sites cobbled together. We generate our C/C++ API reference with Doxygen, our Rust API reference with rustdoc, and everything else with Sphinx. It’s a polyglot docs site. The underlying project supports many programming languages and therefore the docs site needs to wrangle with different documentation generators, because each programming language has its own unique way of generating API reference documentation.
Last month, as I was customizing pigweed.dev to make the cobbled-together
nature of the site less obvious, I got curious about the general state of
polyglot docs sites. Overall it feels like a sparsely explored frontier.
There’s not really a great way to combine the outputs from multiple docs
generators into one cohesive whole yet.
Strategies#
In terms of top-down strategy I can only think of 2 ways to structure a polyglot docs site.
Transformation#
The first strategy is to parse each API reference generator’s output and transform it into markup that plays nicely with your main docs generator. For example, in my first job I ingested Doxygen HTML as input, used XSLT (!!) to transform it into simpler HTML fragments, and then used the raw directive to pull the HTML fragments into my Sphinx site.
The main drawback of the transformation approach boils down to losing out on the expertise of the API reference generators:
Tools like Doxygen, rustdoc, javadoc, etc. understand the details of their respective languages much better than I do. With a custom transformation that “simplifies” the output, there’s a risk that I’m stripping out information that users actually need. I.e. Chesterton’s fence.
These tools have put a lot of thought into the UX of API references. For example, given a structured search query like
vec -> usize, rustdoc’s search engine will only return functions that take in avecas an arg and returnsusize.
Another drawback of transformation is that it goes against the grain of the ecosystem. Rust programmers are familiar with the rustdoc UI. Even if I could theoretically create an API reference that’s superior in every way, I’m still asking my users to figure out a new and different UI that they won’t encounter anywhere else.
Another example of the transformation approach is Breathe. You first run a
Doxygen XML build, and then make that available as an input to the Sphinx
build. In your reStructuredText you insert a directive like
.. doxygenclass:: pw::Foo to indicate the place where the API reference for
pw::Foo should go. Breathe parses the info from the Doxygen XML and
transforms it into API reference content that Sphinx understands. This was the
foundation of C/C++ API reference content on pigweed.dev from 2022 to 2024.
We moved to the approach described in the next section for a few reasons.
Before I explain the issues, please note that I have the utmost respect for the Breathe maintainers and all the people who create and maintain the OSS tools that I depend upon. Breathe just happens to be in the unlucky position of being the transformation tool that I have firsthand experience with.
Now, onto the issues. They all revolve around contributor friction:
One issue was slowness. I don’t remember the exact numbers but Breathe was a significant bottleneck in our docs build. We went from something like 90 seconds with Breathe to 30 seconds without it.
Another issue was too much glue code leading to silent failures. In addition to marking up your headers with Doxygen comments, you have to remember to pull the content into Sphinx via a directive like
.. doxygenclass:: pw::Foo. On quite a few occasions I saw SWEs make an honest effort to document their code, but the documentation never actually got published, because they had forgotten thedoxygenclassstep.The last issue was too much flexibility. Some docs contributors would order their
doxygenclassdirectives alphabetically on a single page. Others would take a thematic approach. E.g. in the middle of a guide on how to foo the bar, they would insert the API reference forpw::Foo.
Turducken#
The second strategy is to defer to the expertise of the API reference
generators and publish their output as-is. This is what pigweed.dev does.
Architecture-wise, it’s a turducken. A chicken stuffed into a duck stuffed
into a turkey. Well, in the case of pigweed.dev it’s more like a chicken
(Doxygen) and a duck (rustdoc) stuffed side-by-side into a turkey (Sphinx).
The obvious problem is that the subsites all look different from each other:
A Doxygen-generated page#
A rustdoc-generated page#
A Sphinx-generated page#
With an obscene amount of !important flags I can make them look more similar to each other. And I plan on doing that. But no amount of CSS hacking can solve the following more insidious problems:
It’s hard to navigate from one subsite to another. E.g. from a rustdoc-generated page there is no way to get back to the pigweed.dev homepage.
The search UX is fragmented and incomplete. E.g. when using the in-site search on a page generated by Sphinx, the search results do not include any content from rustdoc-generated pages.
It’s hard to link from one subsite to another. You end up relying on hardcoded manual relative paths, which are brittle.
So, we’ve got 3 completely separate subsites that have no awareness of each
other’s existence, and we somehow need to make them feel more connected. On
pigweed.dev we made some progress on this front by introducing a universal
header and comprehensive search.
Universal header#
One header to rule them all, one search to find themOne nav to map the path and breadcrumbs to remind them
Every page on the site now has the same top-level links, search, and breadcrumbs UI. The theme selection also stays in sync across the subsites.
The Doxygen-generated page now:
The rustdoc-generated one:
And the Sphinx-generated one:
Implementation-wise, it’s a bunch of postprocessing. During the Doxygen,
rustdoc, and Sphinx builds we inject a <!-- pw-sentinel --> comment into
every page. The Doxygen and rustdoc builds run before the Sphinx build.
We have a Sphinx extension that hooks into the build-finished
event and replaces these comments with the full HTML, CSS, and JS for the
header. The extension has a lot of logic related to constructing the top-level
links and breadcrumbs.
Comprehensive search#
Within the universal header there is now a consistent UI for accessing the in-site search. The search opens as a modal, and results surface as you type. The big win is that the search results are now comprehensive. I.e. all of our Doxygen, rustdoc, and Sphinx content is now indexed and surfaced in results:
The 1st result comes from Doxygen, the 2nd and 3rd come from Sphinx, and the 4th comes from rustdoc.#
The real hero here is Pagefind. This library is a one-stop shop for all of our in-site search needs. We just point Pagefind to the output directory and it creates the search index based off our built HTML. There are both allowlist and denylist APIs for fine-tuning what content gets indexed. For the search UI we use Pagefind’s default web components with a sprinkling of customization. There’s a JS API if you want to completely customize [1] the UX. Pagefind also seems to be doing cool things with web workers, WebAssembly, and on-demand loading of index chunks, but I couldn’t find a good overview of the internals in the official docs.