Wesley Dean
One Documentation Model Across Bash, AWK, Python, and PHP image

One Documentation Model Across Bash, AWK, Python, and PHP

14 min read

Over the last several weeks, I have written about three Doxygen filters: bash-doxygen, awk-doxygen, and python-doxygen. Seen together, that sequence could suggest that I had decided to write a Doxygen filter for every language I use, although that was never the goal. I was trying to create a coherent documentation structure across the languages I use regularly while allowing the maintained source in each language to remain recognizable and appropriate to that language.

That distinction became more important as the work progressed. Bash needed a filter because Doxygen does not understand Bash well enough to infer the structures I wanted documented. AWK needed a different filter because its functions, globals, locals, and pattern/action rules do not map cleanly onto the Bash model. Python already had a mature documentation system, so its problem was less about inventing structure and more about translating Python-native docstrings without allowing Doxygen to become the authoring language. By the time I reached PHP, I had enough examples to ask a better question than “what should the next filter look like?” The more useful question was whether PHP needed a filter at all.

It did not. PHPDoc-style DocBlocks and Doxygen overlap closely enough that I could define a common subset and let Doxygen consume the maintained PHP directly. That result clarified the larger architecture for me: the consistency I wanted belonged in the documentation contract and the generated reference material, not in forcing every source language through identical syntax or identical tooling.

Consistency does not require identical source syntax

One way to pursue consistency would be to require every source file to use the same documentation syntax. I judge that approach to be a poor fit for the languages themselves because a Bash function is not a Python method, an AWK rule is not a PHP class, and each language brings its own syntax, tooling, type model, and expectations about what a maintainer should encounter when opening the source. If the maintained documentation becomes foreign to the language merely so a downstream generator can consume it, the common tooling has begun to dictate too much of the source.

For Bash, I wanted comments that still looked like Bash comments:

##
# Load and validate application configuration.
#
# @param $1 Path to the configuration file.
# @stdout Validated configuration.
# @return 0 on success.
##
load_configuration()
{
  ...
}

bash-doxygen recognizes that documented surface and translates it into a representation Doxygen understands. AWK requires a different treatment because a function may depend on global state, a rule may consume or modify records, and constructs such as getline, subprocesses, and conventional local parameters can matter even when there is no ordinary function-call interface to describe. That is why awk-doxygen solves an AWK-shaped problem rather than copying the Bash filter with a few changed regular expressions.

Python moved the boundary again. It already has docstrings, PEP 257 conventions, type annotations, and established Sphinx/reStructuredText fields:

def load(path: str) -> str:
    """Load and validate a configuration file.

    :param path: Path to the configuration file.
    :returns: The validated configuration.
    :raises ValueError: The configuration is invalid.
    """

Replacing that with Doxygen-specific Python would make the maintained source less Python-like merely to satisfy one publication target, so python-doxygen translates at the Doxygen boundary instead. By that point, three languages had arrived at the same generated documentation system by three different routes, and I no longer saw identical source syntax as a useful measure of consistency.

The common structure is the contract

When I say I want documentation to be consistent across languages, I mean that a maintainer should be able to answer the same kinds of questions even when the syntax used to answer them differs. A documented callable should explain what it does, what information it receives, what each input means, what it produces, how failure becomes visible, what state it reads or modifies, what side effects occur, what assumptions callers must preserve, what security or lifecycle constraints matter, and where durable architectural reasoning can be found when a local implementation depends upon it.

Those questions form a documentation contract that can survive language changes. The source syntax does not have to be identical for the generated reference material to present a coherent model of functions, methods, parameters, return values, exceptions, state, relationships, and behavioral contracts. Doxygen became useful to me as the common publication layer precisely because the languages did not have to agree completely before reaching it. Translation can happen where translation adds value, while native source conventions can remain in place when Doxygen already understands them adequately.

PHP supplied the case that made that distinction concrete. Instead of asking how to translate PHP into Doxygen, I could ask whether a maintained PHPDoc-style DocBlock could already carry the contract I wanted without an intermediate representation.

PHP already speaks enough Doxygen

PHP has a mature documentation convention of its own in PHPDoc-style DocBlocks, and a conventional documented function already looks close to what I would ask someone to write for Doxygen:

/**
 * Load and validate application configuration.
 *
 * Reads the requested configuration document and validates it before returning
 * an application configuration object.  The function does not modify the source
 * file.
 *
 * @param string $path Path to the configuration document.
 * @param bool $strict Whether unknown configuration keys are rejected.
 * @return Configuration Validated configuration owned by the caller.
 * @throws InvalidArgumentException The path or configuration is invalid.
 * @throws RuntimeException The configuration cannot be read.
 */
function loadConfiguration(
    string $path,
    bool $strict = true
): Configuration {
    ...
}

The overlap is semantic as well as visual. PHPDoc and Doxygen both understand a useful common vocabulary that includes:

@param
@return
@throws
@var
@see
@deprecated

Once I reached that point, building another translation layer no longer made architectural sense to me. A filter would have created another implementation to maintain, test, document, release, version, and debug, along with another place where the maintained source and the generated representation could drift apart. If Doxygen can consume a suitable PHPDoc-style DocBlock directly, the lower-risk architecture is to keep that DocBlock as the maintained source and avoid a translation step that does not solve a corresponding problem.

Choosing the overlap deliberately

Direct compatibility still required a standard because two documentation systems that overlap are not necessarily identical. phpDocumentor supports tags beyond the subset I need, while Doxygen has commands that would add little value to ordinary PHP merely because Doxygen recognizes them. I therefore defined the PHP documentation standard around a smaller shared vocabulary: use the common form when it expresses the intended contract accurately, let Doxygen compatibility govern when the systems diverge, and prefer phpDocumentor compatibility when the same maintained source can satisfy both without weakening the Doxygen contract.

That decision affects ordinary prose as well as tags. There is little reason to add Doxygen-specific @brief and @details commands to every PHP DocBlock when PHPDoc/Javadoc-style prose already communicates the same structure. Doxygen can be configured with:

JAVADOC_AUTOBRIEF = YES

With that setting, the first sentence can serve as the brief description and the following prose can provide the details. The maintained PHP remains familiar PHP documentation, while Doxygen receives the structure it needs. The same principle applies to @param, @return, and @throws: those tags already carry useful contracts in both systems, so inventing a second vocabulary would add ceremony without adding information.

There are still cases in which the two systems disagree, and the standard has to own that decision rather than pretending the mismatch does not exist. In this architecture, Doxygen takes precedence because it is the common generated-reference layer across these projects. phpDocumentor compatibility remains useful, though it is subordinate to that explicit publication contract.

Type information can remain language-specific

PHP also reinforced something I had already encountered while building the Python filter: consistency at the documentation-system level does not require the source languages to pretend they share the same type model. Python’s native annotations normally carry the executable type information:

def load(path: str) -> Configuration:
    ...

The Python documentation standard can therefore avoid needlessly duplicating those annotations in ordinary parameter documentation. PHPDoc lives in a different ecosystem. Type information inside @param, @return, and @var is conventional, useful to PHPDoc-aware tooling, and capable of expressing forms that native PHP declarations may not express as conveniently.

The PHP standard therefore permits and expects forms such as:

@param string $path Path to the configuration document.
@return Configuration Validated configuration owned by the caller.

The native declaration remains authoritative for executable behavior, and the DocBlock type must agree with it. I do not need to delete useful PHP-specific type information merely to make the source resemble Python documentation. The cross-language standard needs agreement about the contract being documented and about which representation is authoritative when two representations overlap; it does not require the languages to erase distinctions that are useful in their own ecosystems.

Four languages, one generated documentation system

At this point, the architecture looks less like a collection of filters and more like a set of language-specific paths into one documentation model.

Bash, AWK, and Python source pass through language-specific Doxygen filters while PHP source goes directly to Doxygen; all four produce common reference documentation

Bash and AWK need translation. Python needs selective translation while preserving its native documentation model. PHP can go directly to Doxygen because the PHPDoc/Doxygen common subset gives it a suitable native path. The filters are therefore implementation details inside the larger documentation architecture rather than a goal by themselves.

That gives me a useful criterion when another language enters one of my projects: what is the smallest responsible boundary between that language’s normal documentation practices and the common generated documentation model? Another filter may be appropriate when the source and Doxygen cannot communicate directly. Doxygen’s native support may already be sufficient in other cases, or an established documentation convention may overlap with Doxygen closely enough that a shared subset is the better answer. The existence of the Bash, AWK, and Python filters does not create an obligation to manufacture another one.

Standards matter more than filters

Once the tooling began to converge, I encountered a second problem that the filters themselves could not solve. A filter can describe what syntax it accepts, although that does not tell a developer, reviewer, or coding agent what the project expects the documentation to say. Knowing that @param is valid does not answer whether a parameter description should explain units, ownership, valid ranges, sentinel values, mutation, path interpretation, encoding, or security-sensitive behavior. Knowing that Doxygen supports @throws does not determine which exceptions belong in the caller-visible contract.

The syntax is only part of the standard, so I created coding_standards as the canonical home for reusable coding and documentation standards that can govern multiple projects. The language-specific documentation portion currently looks like this:

standards/
|-- awk/
|   `-- documentation-standard.md
|-- bash/
|   `-- documentation-standard.md
|-- php/
|   `-- documentation-standard.md
`-- python/
    `-- documentation-standard.md

Those files are ordinary Markdown, which matters because they can be read by a person, reviewed as Git changes, supplied to a coding agent before source is modified, and carried into environments where no specialized standards service exists. More importantly, they can describe the documentation contract at a higher level than the parser implementation. The Bash filter does not need to decide what a useful Bash function contract should contain; the Bash documentation standard owns that decision. The Python filter does not need to become a Python type checker; the Python documentation standard can assign semantic validation to Python-native tooling while defining the structured fields that need to cross the Doxygen boundary. The PHP standard can state explicitly that no translation filter belongs in its baseline architecture.

That separation is valuable to me because it keeps syntax recognition, documentation policy, and language semantics in different places. Each piece has a narrower responsibility, and the generated documentation can remain consistent without asking any one tool to understand the entire system.

A standard should be adopted deliberately

Moving the standards into their own repository created one more architectural question: if another project depends upon those standards, how does that project know which revision governs its source? I do not want the answer to be “whatever happens to be on main today.” A repository should be able to carry the exact standards it has adopted, review changes to them, use them without network access during ordinary development, and update them intentionally when a new revision is appropriate.

The coding_standards repository is therefore designed so individual standards can be materialized into a consuming project as ordinary files, for example:

doc/
`-- standards/
    |-- clean-coding-standard.md
    |-- bash-documentation-standard.md
    `-- python-documentation-standard.md

The canonical copy remains in coding_standards; the files inside a consuming repository are pinned snapshots of the standards that project has chosen to adopt. I use bashdeps to manage that relationship. A project can declare an immutable source reference and SHA-256 digest, materialize the exact standard beneath doc/standards/, and then make that local copy available to developers, CI, reviewers, and coding agents as an ordinary repository file.

That arrangement creates an explicit adoption point. A change to the shared standard does not silently change every repository that depends upon it; each project can evaluate the new revision and choose when to update its pinned copy. For documentation policy, I consider that review boundary important because a standard can change expectations for future source changes even when no executable dependency has changed.

The filters were a means to a larger end

I started this work by trying to generate useful Doxygen documentation from Bash, and that effort led to AWK. Python then forced me to distinguish preserving the text of documentation from preserving its structured meaning. PHP supplied the useful counterexample: once a language’s established documentation conventions overlap sufficiently with Doxygen, adding another filter creates machinery without solving a corresponding problem.

The resulting model is the part I intend to keep. Source documentation should fit the language in which it is maintained, language-specific tooling should translate where translation is necessary, Doxygen should provide the common generated reference layer, and the documentation standards should define the contracts people and tools are expected to preserve. The standards now have a home of their own, and individual projects can adopt them as pinned, reviewable dependencies rather than inheriting an ambient rule that changes outside the repository.

Next week, I’ll look at bashdeps: the tool I use to retrieve exact external artifacts, verify them against committed SHA-256 digests, and materialize things such as these standards where a consuming repository expects them.

Tags