Yellow Orange Knowledge Platform

We built Yellow Orange to turn large, permission-sensitive document collections into an AI workspace people can actually trust. It combines OCR, access-aware retrieval, agentic search, and page-level source verification so teams can move from a question to an answer and then back to the evidence behind it.

Why we built it

The knowledge existed. Reaching it was the problem.

The organizations we designed Yellow Orange for did not have a shortage of information. They had policy archives, technical reports, contracts, research papers, meeting packs, and operational manuals. The problem was that the answer to a practical question might be buried on page 73 of a scanned PDF, spread across several documents, or stored in a folder the person asking the question could not access.

Traditional search only helped when someone knew the exact words to look for. A general AI assistant could produce fluent answers, but it did not understand the organization, its permissions, or which document was authoritative. Uploading files to a generic chat also created a new problem: useful answers became disconnected from the governed knowledge system they were supposed to represent.

  • People spent time opening and scanning documents instead of using the knowledge inside them.
  • Teams could not safely search across departmental or project boundaries.
  • Generic AI answers sounded convincing without showing whether the source really supported them.
  • Administrators needed control over users, documents, access, and ongoing operating costs.
Product principle

Trust had to be designed into every layer

We decided early that a good answer was not enough. The user also had to be allowed to see the evidence, the retrieval system had to stay inside that permission boundary, and the interface had to make verification easy. That turned the project from a chatbot into a complete knowledge product.

The resulting architecture treats ingestion, permissions, retrieval, prompting, citations, and the PDF viewer as one chain. If any link is weak, trust breaks. Yellow Orange therefore limits the knowledge available before the model starts reasoning, asks the agent to work from retrieved evidence, and keeps the source attached all the way to the final answer.

  • Permissions are applied before retrieval, not after an answer has been generated.
  • The AI is instructed to search first and to say when the available documents do not contain an answer.
  • Internal citations remain connected to the underlying file and page number.
  • The product supports direct search and conversational research without creating two separate knowledge systems.
Ingestion pipeline

Turning difficult PDFs into usable knowledge

The first technical challenge was getting reliable content out of real organizational documents. Many PDFs have weak text layers, tables, complex layouts, or scanned pages. Yellow Orange stores the private source file, processes it with Mistral OCR, preserves the page relationship, divides the extracted text into focused passages, and creates vector embeddings for retrieval.

Processing happens in the background so an upload can be acknowledged without forcing the user to wait for the complete AI pipeline. Every file has a visible state such as pending, processed, partial, failed, or excluded. That operational detail matters: a knowledge system should show when a source is ready and make failures reviewable instead of silently omitting information.

Onboarding screen explaining file upload and indexing
  • Upload PDFs directly into the workgroups that should own the knowledge.
  • Use OCR to recover content from scanned and layout-heavy documents.
  • Keep page numbers attached to generated passages and embeddings.
  • Track processing and review states so administrators can see what the AI can and cannot use.
Multi-tenant access

The model only sees what the user may see

Yellow Orange separates organizations at the tenant level and divides knowledge again through workgroups. A workgroup can represent a department, vessel, project, client, board, or temporary collaboration. Membership determines which files a person can open and which file identifiers the retrieval tools are allowed to search.

This is a more dependable pattern than asking the model to remember access rules in a prompt. The application calculates the permitted file set from organization and workgroup membership first. That list is then passed into semantic search, direct search, filename lookup, and full-document retrieval. Database policies reinforce the same boundary at the storage layer.

Onboarding screen explaining Yellow Orange workgroups
  • Organization membership defines the primary data boundary.
  • Workgroup membership narrows access to the relevant document collections.
  • Organization and workgroup administrators can invite, approve, and assign users.
  • Role-based application checks and row-level database policies protect the same resources.
Retrieval and reasoning

An agent that researches before it answers

The chat experience is built as a retrieval agent rather than one oversized prompt. It can search for semantically relevant passages, find documents by name, retrieve longer sections when a summary needs more context, and search connected external sources when that feature is enabled. Multiple focused searches can run in parallel and their results are deduplicated before they reach the answer.

The system prompt gives the agent a narrow job: answer from available evidence, cite the sources it uses, ask for clarification when the request is genuinely ambiguous, and do not fill gaps with general model knowledge. The production research behind the platform reinforced this structure. Clear roles, explicit tools, bounded context, citation rules, and defined fallback behavior make an agent more predictable than simply telling it to be accurate.

  • Semantic retrieval finds relevant meaning even when the question uses different wording.
  • Direct search combines vector similarity with keyword signals for a practical search-results view.
  • File discovery and full-document retrieval support summaries and document-specific questions.
  • Optional connected sources and web search extend the workspace without mixing their citation rules.
Verification experience

From answer to evidence in one step

The answer is only the beginning of the workflow. Each internal reference retains the passage, file identity, and page number returned by retrieval. A user can open the original PDF at that page and read the surrounding context before using the information in a decision, report, or conversation.

This verification loop changes how people use AI. Instead of asking colleagues to trust a generated paragraph, the product lets them inspect the evidence themselves. It is especially valuable for policy, legal, technical, and regulated work where the wording around a fact can matter as much as the fact itself.

Onboarding screen showing source-page opening from results
  • Inline references keep claims tied to the passages used to support them.
  • Search results show matching text and the pages where it was found.
  • The PDF viewer opens directly on the cited page.
  • When retrieval finds no supporting material, the assistant is expected to say so.
Operations and governance

A complete product around the AI

A useful internal AI product needs more than retrieval. Yellow Orange includes onboarding, invitations, membership approval, workgroup management, file administration, processing review, conversation history, feedback, and organization-level feature controls. Subscription and seat logic are part of the same operating model.

External source connectors and web search can be enabled per organization instead of appearing as uncontrolled global tools. Administrators decide which sources are available, and the chat route checks both the user entitlement and the organization feature configuration before exposing them to the agent.

  • Manage organizations, roles, workgroups, files, and invitations from dedicated admin flows.
  • Keep conversations and cited sources available for later review.
  • Collect answer feedback to expose where content or retrieval needs improvement.
  • Control premium capabilities and seat access through the billing model.
Why it was innovative

What this build proved for our agency

The innovative part was not one model call. It was the integration of document engineering, permissions, retrieval, agent behavior, and verification into a coherent product. We built a system where the AI cannot casually roam across the whole database, where a scanned page can become searchable, and where the user can challenge an answer by opening its source.

That pattern now informs how Yellow Orange approaches custom AI software. We start with the real decision or workflow, identify the evidence and access boundaries, then design the AI around those constraints. The same architecture can be adapted to technical libraries, compliance archives, contract repositories, research collections, or any domain where a plausible answer is less valuable than a traceable one.

  • AI reliability comes from product architecture as much as from model choice.
  • Permission-aware retrieval is a reusable foundation for enterprise AI applications.
  • Source-first UX gives domain experts a practical way to supervise AI output.
  • The platform demonstrates our ability to take an AI concept through ingestion, application design, security, and operations.

Keep reading

Contact us ->