Auroramind

Aurora Atlas · Documentary corpora for AI systems

Control what feeds the knowledge of your AI systems.

Aurora Atlas turns external sources — public or accessible to your organization — into qualified, structured and continuously maintained documentary corpora.

Find the right sources, select the ones that matter, turn them into usable Markdown resources, and feed Nexus when you want to integrate them into your organization's knowledge.

Assisted ResearchHuman controlContinuous monitoringMarkdown exportNexus preparation

EXTERNAL KNOWLEDGE

StandardsSupplier documentationMarketProfessional sourcesNews
AURORA ATLAS

Documentary corpus

MarkdownNexus-ready

Aurora Atlas

Your AI knows your organization. It also needs to understand its environment.

Your procedures, data, products and internal expertise form the first layer of its knowledge.

But your teams also depend on information that evolves outside the organization: regulations, standards, technologies, supplier documentation, markets, competitors, specialized publications and professional sources your organization is authorized to access.

Atlas helps you select this external knowledge, structure it and keep it current over time.

WHAT YOUR ORGANIZATION ALREADY KNOWS

  • Internal documents
  • Business data
  • Procedures
  • Products and services
  • Team expertise

WHAT IT NEEDS TO FOLLOW EXTERNALLY

  • Regulations and standards
  • Technologies and innovations
  • Manufacturers and suppliers
  • Markets and competitors
  • News and publications
  • Authorized professional sources

Internal knowledge remains your foundation. Atlas adds the external context your AI systems need.

Aurora Atlas

From available information to a controlled documentary corpus

Atlas organizes the full lifecycle of building and maintaining external knowledge.

01

Define

Specify the subject, intent and expected uses of the future corpus.

02

Find

Discover relevant sources or directly add the ones you already know.

03

Qualify

Assess their relevance, quality, depth and identified risks.

04

Build

Collect, clean, deduplicate and transform content into structured documentary resources.

05

Maintain

Monitor selected sources and identify changes that genuinely require new processing.

06

Feed

Retrieve your Markdown datasets or explicitly prepare and send selected resources to Nexus.

Automation accelerates every stage. Your organization remains in control of the documentary scope.

Aurora Atlas

Tell Atlas what you need to know.

You have a subject to document, but not yet a reliable list of sources?

Research starts with your objective, preferred languages and the intended uses of the corpus. Atlas can propose a research strategy and suitable queries before any deep collection begins.

You remain in control of which suggestions are applied and can complement discovery with your own URLs.

  • An explicit documentary intent.
  • Configurable and traceable searches.
  • Sources discovered or manually added.
Aurora Atlas Research interface defining the intent and expected uses of a documentary corpus.
Research starts with the business need and the expected uses of the corpus.

Aurora Atlas

A source being found does not make it a good source.

Collecting more pages does not automatically create better knowledge.

Before deep crawling, Atlas can extract representative content, apply deterministic filters and qualify each source across several dimensions: intrinsic quality, fit with the corpus, usefulness for the intended use cases and depth of content.

The result remains explainable: scores, summaries, verdicts, reasons and risk factors support the proposed decision.

Atlas proposes. You decide.

Sources can be manually selected or deselected before the final corpus is built.

Aurora Atlas Research sources scored, qualified and selected before deep crawling.
Structured scoring and human source selection before deep crawling.

Aurora Atlas

Already know your sources? Collect them directly.

Research is not mandatory.

Atlas can also start from a website, sitemap or known URLs and explore only the relevant areas using inclusion rules, exclusions, depth limits and page limits.

Linked text-based PDFs can be added to the dataset, as can YouTube content when a usable transcript is available.

Public sources

Websites, documentation, help centers, blogs, technical pages and text-based PDFs.

Video sources

YouTube videos when supported subtitles or transcripts are accessible.

Authenticated professional sources

Atlas can reuse a compatible authorized session to access content available to your organization: supplier portals, premium documentation, extranets or professional platforms.

Atlas does not universally bypass CAPTCHAs, paywalls or security mechanisms. Access must be authorized and compatible with the target website.

Aurora Atlas crawler interface with linked PDFs, crawl limits and an authorized authenticated session.
Direct crawling, linked PDFs and authorized authenticated sessions within the same workflow.

Aurora Atlas

Collecting a page is not enough. It needs to become usable.

A web page often contains much more than the information you actually need: navigation, banners, repeated elements, short blocks, consent interfaces or other noise.

Atlas turns collected raw content into cleaner, structured documentary resources.

Mixed sources

Web pages · text PDFs · YouTube · news

Atlas processing

Extraction · cleaning · deduplication · metadata · structuring

Documentary corpus

Individual documents · Markdown dataset · manifest

Markdown is the reference format of the Atlas pipeline. Datasets can be retrieved for your own workflows or continue their journey into Nexus.

Atlas remains useful without Nexus.

Aurora Atlas dataset catalog grouped within a documentary project.
Dataset industrialization within a documentary project.

Aurora Atlas

A good corpus today can become incomplete tomorrow.

A regulation changes. A manufacturer updates its documentation. A competitor publishes a new offer. A technology evolves.

Rebuilding your entire corpus at regular intervals would be costly and generate a large amount of unnecessary processing. Watch provides continuity. Atlas can schedule monitoring of selected sources, detect their evolution and trigger the appropriate pipeline when needed.

Aurora Atlas Watch configuration with scheduling and corpus maintenance pipeline.
Monitoring can extend a Research project through reprocessing and, when configured, all the way to Nexus.

Aurora Atlas

Know what changed before processing everything again.

Atlas distinguishes modified sources from those that remain stable.

Change detection can differentiate meaningful editorial updates from technical or irrelevant variations, helping avoid rebuilding the corpus without a valid reason. An unchanged source can remain untouched. A modified source can re-enter the processing pipeline.

Maintaining knowledge does not mean starting over.

Aurora Atlas source list distinguishing modified and unchanged content.
Modified or unchanged sources: Atlas identifies what genuinely deserves new processing.

Aurora Atlas

Maintaining a corpus also means discovering what did not exist yesterday.

Some knowledge evolves because an existing source changes. Other knowledge evolves because entirely new publications appear.

Atlas currently provides a news feed based on GDELT Cloud v2, linked to Research projects.

Queries remain visible and must be validated before use. New results are deduplicated, qualified within the context of the project and can enter the same processing, Watch and Nexus preparation workflows.

Aurora Atlas GDELT news feed linked to a Research project.
Incremental GDELT news collection integrated into the Research lifecycle.

Aurora Atlas

One mechanism. Very different forms of external knowledge.

01

Regulatory and standards monitoring

Monitor the official publications and reference sources relevant to your business so that teams and AI systems can work with an up-to-date documentary corpus.

Energy regulations, construction standards, sector obligations or professional recommendations.

02

Manufacturer and supplier documentation

Collect and maintain external technical documentation that complements your own procedures and business data.

A support agent can use both internal procedures and up-to-date documentation from the equipment manufacturers used by the organization.

03

Authenticated professional sources

Get more value from information your organization already accesses through subscriptions or contractual services.

A sector platform, extranet or compatible professional knowledge base can become a regularly refreshed source for your AI systems.

04

Technology and market intelligence

Build a corpus around a market, technologies, competitors or trends, then maintain it with new publications and updates from sources already selected.

Innovation monitoring, market intelligence, competitive intelligence or sector observatories.

Aurora Atlas

Add what your organization chooses to follow to what it already knows.

Atlas and Nexus operate at two different stages of the knowledge chain.

INTERNAL ORGANIZATIONAL KNOWLEDGE

  • Documents
  • Données
  • Procédures

EXTERNAL KNOWLEDGE CONTROLLED WITH ATLAS

  • Research
  • Qualification
  • Collection
  • Cleaning
  • Corpus
  • Watch
NEXUS

AI applications and agents

Nexus PocketSilioNexus Knowledge Studiochatbotbusiness agentsother interfaces

AURORA ATLAS

Markdown datasets → your other systems

What your organization knows + what it needs to know about its environment.

ATLAS

Controls the acquisition and preparation of external documentary knowledge.

NEXUS

Then manages ingestion, indexing and exploitation of that knowledge alongside the organization’s other resources.

An AI agent for a construction company can therefore understand its own procedures and products while also using the regulations, manufacturer documentation and innovations that the organization has chosen to follow.

Aurora Atlas

Stay in control through the final step.

Preparing a corpus for Nexus does not automatically send it.

Before transmission, Atlas can apply a Quality Gate combining deterministic rules and, when configured, structured LLM analysis.

Resources can be cleaned, retained, excluded or flagged when uncertainty remains. The user keeps explicit control over the Nexus destination and triggers the transmission.

Prepared

Resources transformed and analyzed.

Sendable

Resources retained after quality control.

Excluded

Resources or chunks removed with an identifiable reason.

Aurora Atlas Quality Gate showing prepared, sendable and excluded sources before Nexus.
Atlas Quality Gate before Nexus ingestion: what enters the knowledge system remains controlled.
Nexus ingestion tracking from Aurora Atlas with completed processing and per-dataset status.
Nexus ingestion tracking

Aurora Atlas

Built to operate in an enterprise environment.

Self-hosted deployment

Atlas can be deployed with Docker on infrastructure controlled by the organization.

User isolation

Data and workflows are scoped to users, with separate user and administrator roles.

Server-side secrets

Provider keys and sensitive credentials are not directly exposed to the browser.

Individual Nexus connections

Each user can have their own Nexus token, stored encrypted server-side.

Controlled external access

Atlas includes protections against unintended network access, with private networks disabled by default for collection.

Explicit destinations

Preparing a corpus and sending it to Nexus remain two separate actions.

Aurora Atlas

Frequently asked questions

Aurora Atlas

Give your AI systems the external knowledge they are missing.

Build a documentary corpus around a market, technology, regulation, your suppliers or professional sources.

Atlas helps you find, qualify, structure and maintain the information you actually choose to bring into your knowledge system.