SourceFusion Free
Build: sourcefusion-storage-2026-07-01-b

Transform unstructured documents with a no-code UI or API.

Use SourceFusion to turn documents, web pages, files, and source text into RAG-ready JSONL, API JSON, extracted tables, metadata, source lineage, and AI-ready records.

No-code processing

Create a SourceFusion job

Upload a file or paste source text. SourceFusion extracts text from text, HTML, PDF, DOCX, PPTX, and XLSX files in the web app; larger source systems and private media routes are provisioned through business deployment.

On by default. Source text and RAG chunk text are processed for the job and not stored in recent-job previews. Usage metadata, file name, page count, status, and billing fields are retained.

DeepWave does not use uploaded documents to train DeepWave models. Turn privacy mode off only for non-sensitive files when you want dashboard previews.

Used only when no file is uploaded or when a file cannot be counted. Uploaded PDFs, Office files, text, HTML, CSV, JSON, XML, and spreadsheets are counted on submit.
No file selected. The manual page count is used only when there is no upload.
Create free account to process
Current estimate

$0.00

1,000requested pages
1,000included pages applied
0billable pages
$0.03/pagestandard page price

DeepWave signal datasets, private connectors, VPC deployment, and customer-specific source domains remain separate product lines.

Platform features

From source intake to AI-ready delivery.

SourceFusion covers the processing path teams expect for RAG, model evaluation, internal search, data products, and governed API delivery.

Business deployments add dedicated instances, private cloud deployment, customer-controlled credentials, source lineage, and processing records for security review.
Hosting and deployment

Run SourceFusion through the hosted app, cloud marketplaces, or private infrastructure.

  • DeepWave hosted processing
  • AWS Marketplace private offer
  • Azure Marketplace private plan
  • Dedicated instance for business accounts
  • Customer VPC deployment for AWS, Azure, or GCP
  • Bare-metal or private endpoint deployment by contract
Extract

Bring source material into one governed processing path.

  • No-code upload
  • API job creation
  • Batch source registration
  • Multi-source configuration
  • Change detection
  • Event-driven updates
  • Incremental processing
  • Connector maintenance for approved sources
Source connectors

Use common enterprise and data-platform sources with customer-controlled credentials.

  • Local upload
  • S3
  • Azure Blob Storage
  • Google Cloud Storage
  • SharePoint
  • OneDrive
  • Google Drive
  • Box
  • Dropbox
  • SFTP
  • GitHub
  • GitLab
  • Confluence
  • Jira
  • Slack
  • Salesforce
  • PostgreSQL
  • MongoDB
  • Snowflake
  • OpenSearch
  • Kafka
  • REST API
Transform

Convert unstructured and semi-structured material into AI-ready records.

  • Document and image extraction
  • Audio/video source registration for private deployments
  • Rich metadata extraction
  • Fast, high-resolution, OCR, and VLM-ready strategies
  • Table extraction
  • Image description enrichment
  • Named-entity extraction
  • Document hierarchy detection
  • Schema evolution
  • Data normalization and flattening
  • Reading-order detection
Supported file intake

Register common enterprise file types for extraction, chunking, enrichment, and delivery.

  • .pdf
  • .docx
  • .pptx
  • .xlsx
  • .csv
  • .tsv
  • .json
  • .xml
  • .html
  • .htm
  • .txt
  • .md
  • .log
  • image and audio/video registration for private extraction workers
Partition

Split documents into useful units before chunking or extraction.

  • Auto strategy
  • Fast strategy
  • High-resolution strategy
  • VLM-ready strategy
  • Page-aware partitioning
  • Table-aware partitioning
  • Private video-to-text route
  • Private speech-to-text route
Chunk

Prepare retrieval units for RAG and downstream model workflows.

  • Chunk by character
  • Chunk by title
  • Chunk by document page
  • Chunk by semantic boundary
  • Contextual chunking
  • Stable chunk IDs
  • Citation metadata
  • Source lineage fields
Enrich

Add context that makes records useful after extraction.

  • Metadata extraction
  • Named-entity recognition
  • OCR enrichment
  • Image description
  • Table description
  • Table-to-HTML output
  • Quality scoring
  • Limitation notes
  • DeepWave signal enrichment for eligible datasets
Embed

Hand off clean chunks to embedding and retrieval systems.

  • Embedding-ready JSONL
  • Azure OpenAI embedding handoff
  • Amazon Bedrock embedding handoff
  • OpenAI embedding handoff
  • Vertex AI handoff
  • IBM watsonx handoff
  • Cohere-compatible handoff
  • Voyage-compatible handoff
  • Customer-key execution for private deployments
Partner integrations

Connect the processing output to the AI platforms teams already use.

  • Amazon Bedrock
  • Azure AI Studio
  • OpenAI
  • Anthropic Claude
  • Gemini
  • Vertex AI
  • IBM watsonx
  • NVIDIA
  • Together AI
Load

Deliver outputs into files, APIs, data warehouses, or retrieval systems.

  • RAG JSONL
  • API JSON
  • CSV
  • Package manifest
  • Source lineage JSON
  • S3 delivery
  • Azure Blob delivery
  • Snowflake delivery
  • PostgreSQL delivery
  • Kafka feed
  • Vector database handoff
Destination connectors

Support downstream data and retrieval platforms.

  • S3
  • Azure Blob Storage
  • Azure AI Search
  • PostgreSQL
  • Snowflake
  • Databricks
  • Elasticsearch
  • OpenSearch
  • MongoDB
  • Neo4j
  • Pinecone
  • Qdrant
  • Redis
  • Weaviate
  • LanceDB
  • DuckDB
Orchestration

Operate repeatable processing jobs instead of one-off scripts.

  • Full ETL orchestration
  • Smart document routing
  • Workflow scheduling
  • Workflow optimization
  • Role-based access control
  • Retry handling
  • Error transparency
  • Duplicate prevention
  • Metadata propagation
  • Custom plugins for private deployments
Security

Protect customer inputs and processing credentials.

  • Privacy mode for web processing
  • No DeepWave model-training use of customer uploads
  • User-controlled purge for stored previews
  • Customer-controlled credentials
  • Encrypted transport
  • Permission-based access
  • Private deployment option
  • Dedicated instance option
  • Source manifests
  • Processing audit records
  • Zero raw card data stored by DeepWave
Compliance and support

Support enterprise review and procurement.

  • AWS and Azure procurement paths
  • Data-processing manifests
  • Source lineage exports
  • Support channel for business plans
  • Security questionnaire support
  • Private deployment review
  • Renewal and entitlement records
Security and deployment

Run the processing path where the data belongs.

SourceFusion is designed for customer-controlled credentials, encrypted transport, permissioned access, source-level manifests, and customer-controlled retention. Enterprise deployments can be delivered as hosted SaaS, a dedicated instance, a private marketplace plan, or a customer VPC deployment.

Privacy

No training use

Uploaded customer documents are processed for the customer job and are not used to train DeepWave models.

Retention

Privacy mode and purge

Privacy mode avoids stored source/chunk previews. Stored previews can be purged from the job record by the account owner.

Deployment

Hosted, dedicated, or VPC

Business accounts can use DeepWave hosted processing, private AWS/Azure procurement, or customer cloud deployment.

API

Estimate and create jobs programmatically.

The estimate endpoint is public for pricing previews. Job creation requires an authenticated account so page allowance and usage can be tracked.

POST /api/sourcefusion/estimate
{
  "sourceName": "Policy archive",
  "sourceType": "PDF / report",
  "requestedPages": 18000,
  "outputFormat": "RAG JSONL + API JSON",
  "privacyMode": true
}

POST /api/sourcefusion/jobs
Authorization: DeepWave account session
{
  "sourceName": "Policy archive",
  "sourceType": "PDF / report",
  "requestedPages": 18000,
  "outputFormat": "RAG JSONL + API JSON",
  "privacyMode": true,
  "sourceText": "Optional source text for immediate chunking"
}