Platform features
From source intake to AI-ready delivery.
SourceFusion covers the processing path teams expect for RAG, model evaluation, internal search, data products, and governed API delivery.
Business deployments add dedicated instances, private cloud deployment, customer-controlled credentials, source lineage, and processing records for security review.
Hosting and deployment
Run SourceFusion through the hosted app, cloud marketplaces, or private infrastructure.
- DeepWave hosted processing
- AWS Marketplace private offer
- Azure Marketplace private plan
- Dedicated instance for business accounts
- Customer VPC deployment for AWS, Azure, or GCP
- Bare-metal or private endpoint deployment by contract
Extract
Bring source material into one governed processing path.
- No-code upload
- API job creation
- Batch source registration
- Multi-source configuration
- Change detection
- Event-driven updates
- Incremental processing
- Connector maintenance for approved sources
Source connectors
Use common enterprise and data-platform sources with customer-controlled credentials.
- Local upload
- S3
- Azure Blob Storage
- Google Cloud Storage
- SharePoint
- OneDrive
- Google Drive
- Box
- Dropbox
- SFTP
- GitHub
- GitLab
- Confluence
- Jira
- Slack
- Salesforce
- PostgreSQL
- MongoDB
- Snowflake
- OpenSearch
- Kafka
- REST API
Transform
Convert unstructured and semi-structured material into AI-ready records.
- Document and image extraction
- Audio/video source registration for private deployments
- Rich metadata extraction
- Fast, high-resolution, OCR, and VLM-ready strategies
- Table extraction
- Image description enrichment
- Named-entity extraction
- Document hierarchy detection
- Schema evolution
- Data normalization and flattening
- Reading-order detection
Supported file intake
Register common enterprise file types for extraction, chunking, enrichment, and delivery.
- .pdf
- .docx
- .pptx
- .xlsx
- .csv
- .tsv
- .json
- .xml
- .html
- .htm
- .txt
- .md
- .log
- image and audio/video registration for private extraction workers
Partition
Split documents into useful units before chunking or extraction.
- Auto strategy
- Fast strategy
- High-resolution strategy
- VLM-ready strategy
- Page-aware partitioning
- Table-aware partitioning
- Private video-to-text route
- Private speech-to-text route
Chunk
Prepare retrieval units for RAG and downstream model workflows.
- Chunk by character
- Chunk by title
- Chunk by document page
- Chunk by semantic boundary
- Contextual chunking
- Stable chunk IDs
- Citation metadata
- Source lineage fields
Enrich
Add context that makes records useful after extraction.
- Metadata extraction
- Named-entity recognition
- OCR enrichment
- Image description
- Table description
- Table-to-HTML output
- Quality scoring
- Limitation notes
- DeepWave signal enrichment for eligible datasets
Embed
Hand off clean chunks to embedding and retrieval systems.
- Embedding-ready JSONL
- Azure OpenAI embedding handoff
- Amazon Bedrock embedding handoff
- OpenAI embedding handoff
- Vertex AI handoff
- IBM watsonx handoff
- Cohere-compatible handoff
- Voyage-compatible handoff
- Customer-key execution for private deployments
Partner integrations
Connect the processing output to the AI platforms teams already use.
- Amazon Bedrock
- Azure AI Studio
- OpenAI
- Anthropic Claude
- Gemini
- Vertex AI
- IBM watsonx
- NVIDIA
- Together AI
Load
Deliver outputs into files, APIs, data warehouses, or retrieval systems.
- RAG JSONL
- API JSON
- CSV
- Package manifest
- Source lineage JSON
- S3 delivery
- Azure Blob delivery
- Snowflake delivery
- PostgreSQL delivery
- Kafka feed
- Vector database handoff
Destination connectors
Support downstream data and retrieval platforms.
- S3
- Azure Blob Storage
- Azure AI Search
- PostgreSQL
- Snowflake
- Databricks
- Elasticsearch
- OpenSearch
- MongoDB
- Neo4j
- Pinecone
- Qdrant
- Redis
- Weaviate
- LanceDB
- DuckDB
Orchestration
Operate repeatable processing jobs instead of one-off scripts.
- Full ETL orchestration
- Smart document routing
- Workflow scheduling
- Workflow optimization
- Role-based access control
- Retry handling
- Error transparency
- Duplicate prevention
- Metadata propagation
- Custom plugins for private deployments
Security
Protect customer inputs and processing credentials.
- Privacy mode for web processing
- No DeepWave model-training use of customer uploads
- User-controlled purge for stored previews
- Customer-controlled credentials
- Encrypted transport
- Permission-based access
- Private deployment option
- Dedicated instance option
- Source manifests
- Processing audit records
- Zero raw card data stored by DeepWave
Compliance and support
Support enterprise review and procurement.
- AWS and Azure procurement paths
- Data-processing manifests
- Source lineage exports
- Support channel for business plans
- Security questionnaire support
- Private deployment review
- Renewal and entitlement records