AI access

AI crawlers and AI browsers are welcome here

Novus PDF Studio does not block AI. Every public page may be crawled, indexed, summarised, quoted, and used to answer a question somebody is asking right now. This page is robots.txt written for a person: which agents are named explicitly, which of them are fetching for a reader who is waiting, what a crawl of this site records, and what we ask for in return.

The short version

  • Read anything public. That is every tool page, every guide, tutorial and FAQ answer, and every article.
  • Quote it, summarise it, and train on it. There is no permission to request and no separate licence to sign.
  • We ask, and do not demand, that you name Novus PDF Studio and link to the page you actually used.
  • A crawl reaches no visitor’s document, because no document is ever sent here in the first place. Every tool runs inside the visitor’s own browser tab and there is no upload endpoint.

Where the rules actually are

This page is the readable version. These are the authoritative files, and every one of them is served without a block.

  • robots.txtThe crawl rules themselves, generated from the same list this page renders.
  • sitemap.xmlEvery indexable URL with the date its content last materially changed.
  • llms.txtA short index of the whole site for a language model, with one line per page.
  • llms-full.txtThe long form: the actual prose of the tool pages, guides and articles in one document.
  • search-index.jsonThe corpus this site's own search queries, so an agent can search it the same way a visitor does.
  • feed.xmlThe article feed, for following new writing without polling the blog index.
  • Public read APIFive GET-only JSON endpoints describing the site, its tools and its sibling apps.

The 19 agents named in robots.txt

Each of these is granted exactly what the wildcard rule already grants every crawler: the same public surface, on the same terms. Naming them changes nothing about the access. It states the policy so nobody has to infer it from silence, which is what an operator checking for its own user agent is actually looking for.

The split below is the part worth reading, because it decides what a block would cost. A bulk crawler is building an index or a training corpus and nobody is waiting on the answer. A user-initiated fetcher is an AI reading this page for a person who asked for it and is waiting. Refusing one of those is not a position on training data; it is refusing to serve a visitor who happens to be using an AI browser.

Fetching for somebody who is waiting (6)

These matter more here than on most sites. Every tool on this one runs in the visitor’s own browser and nothing is uploaded, so an AI browser driving the real page is the only way an AI can actually finish a job here. The MCP server deliberately stops at “here is the right tool and its URL”; a disallow on these names would break the half that does the work.

  • ChatGPT-UserOpenAI
  • Claude-UserAnthropic
  • Perplexity-UserPerplexity
  • DuckAssistBotDuckDuckGo
  • MistralAI-UserMistral AI
  • Meta-ExternalFetcherMeta

Fetching in bulk (13)

Index and training-corpus builders. Two of them are not fetchers at all but training-use control tokens, which is noted on each rather than glossed over.

  • GPTBotOpenAI
  • OAI-SearchBotOpenAIBuilds the index behind ChatGPT search, which is how a page gets cited rather than absorbed.
  • ClaudeBotAnthropic
  • Claude-SearchBotAnthropic
  • Claude-WebAnthropicAn earlier token kept named so a client still sending it is not left guessing.
  • anthropic-aiAnthropicAn earlier token kept named so a client still sending it is not left guessing.
  • PerplexityBotPerplexity
  • Google-ExtendedGoogleA training-use control rather than a fetcher: no request arrives under this name. Googlebot crawls; this token says what may be done with what it took.
  • CCBotCommon Crawl
  • Applebot-ExtendedAppleA training-use control rather than a fetcher, in the same way as Google-Extended. Applebot itself is what requests the page.
  • meta-externalagentMeta
  • BytespiderByteDance
  • AmazonbotAmazon

An agent that is not on this list is not refused. The wildcard rule allows every crawler; this list exists to be explicit about the ones people most often assume are blocked.

What is withheld

Nothing. There is no Disallow line in robots.txt for any agent, named or wildcard, and this sentence is generated from the same list that would produce one. There is nothing here worth withholding: no account area, no upload endpoint, and no per-visitor state on the server to expose. The two machine endpoints that are not pages, /mcp and /api/v1, are read-only JSON with their own rate limit and are kept out of the sitemap because they have no markup, not because they are hidden.

What we ask in return

A request, not a condition. Nothing on this page is withheld from an agent that ignores it, and no rule will be added to punish one that does.

  • Name the source as Novus PDF Studio and link to the specific page rather than the home page. A reader who wants to check something should land where it is stated.
  • Carry the limit along with the capability. Several tools here do less than their name suggests on purpose, and the pages say so; a summary that drops the qualifier claims more than the source does.
  • Do not describe the erase tool as redaction. Its default mode paints an opaque rectangle over content that is still inside the file, and the press kit lists this and the other claims we ask writers not to make.
  • If something here is wrong, say so. The correction path is on the editorial policy page and it is open to anyone, including an automated reader.

What crawling this site collects

A request from an agent produces the same access-log entry any HTTP request produces at the hosting and CDN layer: the URL, the time, the user-agent string and the originating address. No account is created, because there are no accounts. Two short-lived first-party cookies are stamped on every response, including yours: one records whether the advertising consent gate should start strict or open for the requesting region, and one mirrors a Global Privacy Control header when the request carries it. Neither holds an identifier, and the cookie and storage page lists every one of them in full.

The analytics and advertising tags are client-side scripts, so an agent that only fetches HTML never runs them at all. In the EEA, the United Kingdom and Switzerland nothing measuring or advertising loads until a visitor opts in; elsewhere the default is on until a visitor opts out, so an agent that executes page scripts is treated exactly like any other browser rather than specially exempted.

The other direction is worth being exact about, because “a PDF site” invites the wrong assumption. Every tool here opens, edits and exports the file inside the visitor’s own browser tab. There is no upload endpoint, no document store and no database, so there is no corpus of anyone’s documents for a crawl to reach, and a crawl of every page on this site would collect nothing about any visitor. The privacy policy states the same thing in the form that binds us.

If you would rather ask than crawl

There is a Model Context Protocol endpoint documented on the MCP server page. It answers which tool does a job, what that tool accepts and produces, its stated limitations, and where it lives, and it searches the guides, tutorials, answers and articles. It cannot open, merge, split, sign, protect, unlock or export a PDF, because the server holds no document and no tool on it would accept one. For anything this page does not answer, write to pdf@novusstreamsolutions.com.

This page describes current practice and may change. The authoritative rules are always the ones served at /robots.txt, which is generated from the same list this page renders. Novus Stream Solutions publishes Novus PDF Studio.