> ## Documentation Index
> Fetch the complete documentation index at: https://docs.verbex.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Discover a Website's URLs

> Discovers the URLs published by a website, so you can choose which to crawl.

This is a lookup helper and is not scoped to a knowledge base - note the absence of kb_id in the path. Nothing is ingested; feed the URLs you want into the website crawl endpoint.



## OpenAPI

````yaml /api-reference/openapi.json get /api/v1/knowledge-bases/documents/website/sitemaps
openapi: 3.1.0
info:
  title: Verbex Platform API
  description: API for managing AI agents, calls, phone numbers, and more.
  version: 1.0.0
servers: []
security: []
tags:
  - name: Knowledge Bases
    description: >-
      Create and manage knowledge bases, ingest documents into them, and search
      them semantically.


      **Authentication.** Every request carries a bearer token: `Authorization:
      Bearer <API_TOKEN>`. That is the only credential a client sends. The
      gateway resolves your identity from it and injects the tenant headers the
      service reads internally, so clients never send `x-user-org-id` or
      `x-verbex-id` themselves.


      **Tenancy is enforced on every route.** A knowledge base belongs to one
      organization and one user. Requesting one your token does not own returns
      `404 Knowledge base not found`, never `403`. This is deliberate: a
      nonexistent ID, a malformed ID and someone else's ID are indistinguishable
      in the response, so the API cannot be used to probe for other tenants'
      data. If you get a `404` on an ID you are certain exists, suspect the
      token before you suspect the ID.


      **Ingestion is asynchronous.** File upload and website crawl both answer
      `202` as soon as the material is stored and the job is queued. Parsing,
      chunking, embedding and indexing happen afterwards in a worker. Poll the
      document status endpoint until it reports `completed` before expecting
      search to see the content. That gap is the most common source of "why
      isn't my data showing up".


      **There is no document update endpoint.** Every submission mints a new
      `document_id`, so re-uploading a file or re-crawling the same URLs creates
      a second document while the first remains. To refresh material, delete
      then resubmit — and note that the sequence is not atomic.
paths:
  /api/v1/knowledge-bases/documents/website/sitemaps:
    get:
      tags:
        - Knowledge Bases
      summary: Discover a Website's URLs
      description: >-
        Discovers the URLs published by a website, so you can choose which to
        crawl.


        This is a lookup helper and is not scoped to a knowledge base - note the
        absence of kb_id in the path. Nothing is ingested; feed the URLs you
        want into the website crawl endpoint.
      operationId: >-
        fetch_website_sitemaps_api_v1_knowledge_bases_documents_website_sitemaps_get
      parameters:
        - name: url
          in: query
          required: true
          schema:
            type: string
            title: Url
          description: Base URL of the website to inspect.
      responses:
        '200':
          description: URLs discovered for the site.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/SiteMapResponse'
        '401':
          description: Missing or invalid bearer token.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/KnowledgeBaseErrorResponse'
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
        '503':
          description: >-
            Service unavailable. A database failure, not a problem with your
            payload.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/KnowledgeBaseErrorResponse'
components:
  schemas:
    SiteMapResponse:
      properties:
        success:
          type: boolean
          title: Success
          description: Whether the site map was successfully retrieved.
        links:
          items:
            type: string
          type: array
          title: Links
          description: >-
            URLs discovered on the site. Feed the ones you want into the website
            crawl endpoint.
      type: object
      required:
        - success
        - links
      title: SiteMapResponse
      description: URLs discovered for a website.
    KnowledgeBaseErrorResponse:
      properties:
        error:
          type: string
          title: Error
          description: A short error code identifying the type of error that occurred.
        message:
          type: string
          title: Message
          description: >-
            A detailed human-readable message explaining the error and possible
            solutions.
      type: object
      required:
        - error
        - message
      title: KnowledgeBaseErrorResponse
    HTTPValidationError:
      properties:
        detail:
          items:
            $ref: '#/components/schemas/ValidationError'
          type: array
          title: Detail
      type: object
      title: HTTPValidationError
    ValidationError:
      properties:
        loc:
          items:
            anyOf:
              - type: string
              - type: integer
          type: array
          title: Location
        msg:
          type: string
          title: Message
        type:
          type: string
          title: Error Type
      type: object
      required:
        - loc
        - msg
        - type
      title: ValidationError

````