EmbedCrate: self-hosted text embeddings and reranking server powered by Hugging Face TEI.
10K+
Open-source, self-hosted text embeddings and reranking API.
GitHub: https://github.com/hwdsl2/embedcrateβ
Run text embeddings and reranking on your own server with EmbedCrate. Powered by Hugging Face Text Embeddings Inference (TEI)β , it provides an OpenAI-compatible /v1/embeddings API and an optional /rerank endpoint, with configurable models and persistent caching.
Previously known as
docker-embeddings, maintained by hwdsl2β . The Docker image remainshwdsl2/embeddings-server.
Features:
POST /v1/embeddings for embedding requests from compatible OpenAI SDKs and apps.POST /rerank with a cross-encoder model to re-score retrieved documents for higher retrieval accuracy.BAAI/bge-small-en-v1.5, BAAI/bge-m3, nomic-embed-text-v1.5, and more.EMBED_LOCAL_ONLY).Also available as part of the Self-Hosted AI Stackβ , which deploys a complete self-hosted AI stack with a single command.
π The Self-Hosted AI Builderβs Guideβ is a practical guide to building, securing, and operating your own private AI stack.
Also available:
Use this command to set up a text embeddings server:
docker run \
--name embeddings \
--restart=always \
-v embeddings-data:/var/lib/embeddings \
-p 8000:8000 \
-d hwdsl2/embeddings-server
Note: For internet-facing deployments, use a reverse proxyβ to add HTTPS. Also replace -p 8000:8000 with -p 127.0.0.1:8000:8000 in the docker run command above, to prevent direct access to the unencrypted port.
The default model BAAI/bge-small-en-v1.5 (~130 MB) is downloaded and cached on first start. Check the logs to confirm the server is ready:
docker logs embeddings
Once you see "Text embeddings server is ready", generate your first embeddings:
Fresh persistent installations require an API key. Retrieve it for the following examples:
embed_api_key="$(docker exec embeddings embed_manage --getkey)"
curl http://your_server_ip:8000/v1/embeddings \
-H "Authorization: Bearer $embed_api_key" \
-H "Content-Type: application/json" \
-d '{"input": "The quick brown fox", "model": "text-embedding-ada-002"}'
Response:
{"object":"list","data":[{"object":"embedding","embedding":[0.032,...,-0.017],"index":0}],"model":"BAAI/bge-small-en-v1.5","usage":{"prompt_tokens":5,"total_tokens":5}}
amd64 (x86_64), arm64 (aarch64, e.g. AWS Graviton, Apple Silicon VMs)BAAI/bge-small-en-v1.5 model (see model tableβ )EMBED_LOCAL_ONLY=true with pre-cached models.For internet-facing deployments, see Using a reverse proxyβ to add HTTPS.
Get the trusted build from the Docker Hub registryβ :
docker pull hwdsl2/embeddings-server
Alternatively, you may download from Quay.ioβ :
docker pull quay.io/hwdsl2/embeddings-server
docker image tag quay.io/hwdsl2/embeddings-server hwdsl2/embeddings-server
Supported platforms: linux/amd64, linux/arm64.
All variables are optional. Fresh installs with a mounted /var/lib/embeddings volume auto-generate a Bearer token. Existing installs without a key remain open for backward compatibility.
This Docker image uses the following variables, that can be declared in an env file (see exampleβ ):
| Variable | Description | Default |
|---|---|---|
EMBED_MODEL | HuggingFace model ID to use for embeddings. See model tableβ for options. | BAAI/bge-small-en-v1.5 |
EMBED_PORT | HTTP port for the API (1β65535). | 8000 |
EMBED_API_KEY | Optional Bearer token. Fresh persistent installs auto-generate one. If set, all API requests must include Authorization: Bearer <key>. Set explicitly empty to disable authentication. | Auto-generated for fresh persistent installs |
EMBED_HF_TOKEN | HuggingFace Hub token for accessing private or gated models. Not required for public models. | (not set) |
EMBED_LOCAL_ONLY | When set to any non-empty value (e.g. true), disables all HuggingFace model downloads. For offline or air-gapped deployments with pre-cached models. | (not set) |
EMBED_ENABLED | Set to false to disable the embeddings process (for rerank-only mode). | true |
RERANK_ENABLED | Set to true to enable the reranking server (cross-encoder model on a separate port). | (not set) |
RERANK_MODEL | HuggingFace cross-encoder model ID for reranking. See reranker modelsβ . | BAAI/bge-reranker-v2-m3 |
RERANK_PORT | HTTP port for the reranker API. Defaults to 8000 if embeddings is disabled. | 8001 |
RERANK_API_KEY | Optional Bearer token for the reranker. Falls back to EMBED_API_KEY if unset. Set explicitly empty to disable reranker authentication. | (falls back to EMBED_API_KEY) |
EMBED_DISABLE_USAGE_COUNTS | Set to 1 to disable anonymous aggregate usage counts. | (not set) |
Note: In your env file, you may enclose values in single quotes, e.g. VAR='value'. Do not add spaces around =. If you change EMBED_PORT, update the -p flag in the docker run command accordingly.
Example using an env file:
cp embed.env.example embed.env
# Edit embed.env with your settings, then:
docker run \
--name embeddings \
--restart=always \
-v embeddings-data:/var/lib/embeddings \
-v ./embed.env:/embed.env:ro \
-p 8000:8000 \
-d hwdsl2/embeddings-server
The env file is bind-mounted into the container, so changes are picked up on every restart without recreating the container.
--env-filedocker run \
--name embeddings \
--restart=always \
-v embeddings-data:/var/lib/embeddings \
-p 8000:8000 \
--env-file=embed.env \
-d hwdsl2/embeddings-server
cp embed.env.example embed.env
# Edit embed.env as needed, then:
docker compose up -d
docker logs embeddings
Example docker-compose.yml (already included):
services:
embeddings:
image: hwdsl2/embeddings-server
container_name: embeddings
restart: always
ports:
- "8000:8000/tcp" # For a host-based reverse proxy, change to "127.0.0.1:8000:8000/tcp"
# - "8001:8001/tcp" # Reranker API (uncomment if RERANK_ENABLED=true in embed.env)
volumes:
- embeddings-data:/var/lib/embeddings
- ./embed.env:/embed.env:ro
volumes:
embeddings-data:
name: embeddings-data
Note: For internet-facing deployments, use a reverse proxyβ to add HTTPS. Also change "8000:8000/tcp" to "127.0.0.1:8000:8000/tcp" in docker-compose.yml, to prevent direct access to the unencrypted port.
The API is compatible with OpenAI's embeddings endpointβ . For clients using the OpenAI SDK, configure the base URL and your server's API key:
The /v1/embeddings endpoint is served directly by TEI. Supported OpenAI request fields depend on TEI; fields such as encoding_format, dimensions, user, and token-array inputs are upstream-dependent and not documented or tested by this image.
Fresh persistent installations require an API key. Retrieve it for the following examples:
embed_api_key="$(docker exec embeddings embed_manage --getkey)"
export OPENAI_BASE_URL="http://your_server_ip:8000"
export OPENAI_API_KEY="$embed_api_key"
If API key authentication is disabled, omit the Authorization header in curl examples. OpenAI SDK clients still require a nonempty key; set OPENAI_API_KEY=unused.
POST /v1/embeddings
Content-Type: application/json
Parameters:
| Parameter | Type | Required | Description |
|---|---|---|---|
input | string or array | β | Text to embed. Pass a string for a single input or an array of strings for batch embedding. |
model | string | β | Pass any string (e.g. text-embedding-ada-002). The value is accepted for API compatibility; the active model set by EMBED_MODEL is always used. |
Example β single input:
curl http://your_server_ip:8000/v1/embeddings \
-H "Authorization: Bearer $embed_api_key" \
-H "Content-Type: application/json" \
-d '{"input": "The quick brown fox", "model": "text-embedding-ada-002"}'
Example β batch input:
curl http://your_server_ip:8000/v1/embeddings \
-H "Authorization: Bearer $embed_api_key" \
-H "Content-Type: application/json" \
-d '{"input": ["First sentence", "Second sentence"], "model": "text-embedding-ada-002"}'
With API key authentication:
curl http://your_server_ip:8000/v1/embeddings \
-H "Authorization: Bearer $embed_api_key" \
-H "Content-Type: application/json" \
-d '{"input": "Your text here", "model": "text-embedding-ada-002"}'
Response:
{
"object": "list",
"data": [
{
"object": "embedding",
"embedding": [0.032, -0.018, ...],
"index": 0
}
],
"model": "BAAI/bge-small-en-v1.5",
"usage": { "prompt_tokens": 5, "total_tokens": 5 }
}
GET /info
Returns the active model ID, maximum input length, and server version.
curl http://your_server_ip:8000/info \
-H "Authorization: Bearer $embed_api_key"
Requires
RERANK_ENABLED=truein your env file. The reranker runs on port 8001 by default.
POST /rerank
Content-Type: application/json
Parameters:
| Parameter | Type | Required | Description |
|---|---|---|---|
query | string | β | The search query to rank documents against. |
texts | array of strings | β | The documents to rerank. |
raw_scores | boolean | If true, returns raw cross-encoder scores instead of normalized scores. Default: false. | |
truncate | boolean | If true, truncates inputs that exceed the model's max length. Default: true. |
Example:
By default, the reranker uses the embeddings key. If you configured a separate RERANK_API_KEY, replace the value below with that key. If reranker authentication is disabled, omit the Authorization header.
rerank_api_key="$embed_api_key"
curl http://your_server_ip:8001/rerank \
-H "Authorization: Bearer $rerank_api_key" \
-H "Content-Type: application/json" \
-d '{
"query": "What is deep learning?",
"texts": [
"Deep learning is a subset of machine learning...",
"The weather today is sunny with a high of 75Β°F.",
"Neural networks are inspired by the human brain."
],
"raw_scores": false
}'
Response:
[
{"index": 0, "score": 0.98},
{"index": 2, "score": 0.72},
{"index": 1, "score": 0.01}
]
Results are sorted by relevance score (highest first). Use this to re-rank documents retrieved by embeddings similarity search.
An interactive Swagger UI is available at:
http://your_server_ip:8000/docs
If reranking is enabled, the reranker also has its own interactive docs at:
http://your_server_ip:8001/docs
All server data is stored in the Docker volume (/var/lib/embeddings inside the container):
/var/lib/embeddings/
βββ models--BAAI--bge-small-en-v1.5/ # Cached embedding model files
βββ models--BAAI--bge-reranker-v2-m3/ # Cached reranker model files (if enabled)
βββ .port # Active port (used by embed_manage)
βββ .model # Active model ID (used by embed_manage)
βββ .rerank_model # Active reranker model (used by embed_manage)
βββ .rerank_port # Active reranker port (used by embed_manage)
βββ .server_addr # Cached server IP (used by embed_manage)
Back up the Docker volume to preserve downloaded models. Models range from ~90 MB to ~1.3 GB and are only downloaded once; preserving the volume avoids re-downloading on container recreation.
Use embed_manage inside the running container to inspect and manage the server.
Show server info:
docker exec embeddings embed_manage --showinfo
List recommended models:
docker exec embeddings embed_manage --listmodels
List recommended reranker models:
docker exec embeddings embed_manage --listrerankers
Pre-download a model:
docker exec embeddings embed_manage --pullmodel BAAI/bge-base-en-v1.5
docker exec embeddings embed_manage --pullmodel BAAI/bge-reranker-v2-m3
To change the active model:
(Optional but recommended) Pre-download the new model while the server is running:
docker exec embeddings embed_manage --pullmodel BAAI/bge-base-en-v1.5
Update EMBED_MODEL in your embed.env file (or add -e EMBED_MODEL=BAAI/bge-base-en-v1.5 to your docker run command).
Restart the container:
docker restart embeddings
Recommended models:
| Model | Disk | RAM (approx) | Notes |
|---|---|---|---|
BAAI/bge-small-en-v1.5 | ~130 MB | ~250 MB | Fastest; English β default |
BAAI/bge-base-en-v1.5 | ~440 MB | ~700 MB | Good balance; English |
BAAI/bge-large-en-v1.5 | ~1.3 GB | ~2 GB | High accuracy; English |
BAAI/bge-m3 | ~570 MB | ~1 GB | Multilingual; cross-lingual retrieval |
nomic-ai/nomic-embed-text-v1.5 | ~550 MB | ~1 GB | Multilingual; long context (8192 tokens) |
sentence-transformers/all-MiniLM-L6-v2 | ~90 MB | ~200 MB | Very small; fast; popular for semantic search |
Tip:
BAAI/bge-m3andnomic-ai/nomic-embed-text-v1.5are recommended for non-English or multilingual workloads. For English RAG pipelines,BAAI/bge-base-en-v1.5offers a good accuracy-to-resource balance.
Models are cached in the /var/lib/embeddings Docker volume and only downloaded once. Any HuggingFace model supported by TEI can be used β see the TEI supported models listβ .
Reranking improves retrieval quality by re-scoring documents with a cross-encoder model. Enable it by setting RERANK_ENABLED=true in your env file.
Add to your embed.env:
RERANK_ENABLED=true
Expose port 8001 (add -p 8001:8001 to your docker run command, or uncomment the port in docker-compose.yml).
Restart the container:
docker restart embeddings
The reranker model (BAAI/bge-reranker-v2-m3, ~560 MB) is downloaded on first start.
| Mode | Configuration | Memory (approx) |
|---|---|---|
| Embeddings only (default) | RERANK_ENABLED unset | ~250 MB (bge-small) |
| Embeddings + Reranking | RERANK_ENABLED=true | ~850 MB (bge-small + bge-reranker-v2-m3) |
| Reranking only | EMBED_ENABLED=false, RERANK_ENABLED=true | ~600 MB (bge-reranker-v2-m3) |
In rerank-only mode, the reranker listens on port 8000 by default (since the embeddings process is disabled), unless RERANK_PORT is explicitly set.
| Model | Disk | RAM (approx) | Notes |
|---|---|---|---|
BAAI/bge-reranker-v2-m3 | ~560 MB | ~600 MB | Multilingual; strong accuracy β default |
BAAI/bge-reranker-base | ~440 MB | ~500 MB | English; good balance |
BAAI/bge-reranker-large | ~1.3 GB | ~1.5 GB | English; highest accuracy |
cross-encoder/ms-marco-MiniLM-L6-v2 | ~90 MB | ~150 MB | Very small; fast; English |
To use the reranker with LiteLLM, including GatewayCrateβ , add it as a rerank model in your LiteLLM config:
model_list:
- model_name: rerank
litellm_params:
model: huggingface/BAAI/bge-reranker-v2-m3
api_base: http://embeddings:8001
api_key: os.environ/RERANK_API_KEY
Set RERANK_API_KEY in the LiteLLM container environment to the key accepted by your reranker. By default, this is the embeddings key retrieved above; use your separately configured reranker key if applicable. If reranker authentication is disabled, omit the api_key entry.
Then call the LiteLLM /rerank endpoint, and it will proxy to your self-hosted reranker.
See Securing your serverβ .
For internet-facing deployments, place a reverse proxy in front of the embeddings server to handle HTTPS termination. The server works without HTTPS on a local or trusted network, but HTTPS is recommended when the API endpoint is exposed to the internet.
Use one of the following addresses to reach the embeddings container from your reverse proxy:
embeddings:8000 β if your reverse proxy runs as a container in the same Docker network as the embeddings server (e.g. defined in the same docker-compose.yml).127.0.0.1:8000 β if your reverse proxy runs on the host and port 8000 is published (the default docker-compose.yml publishes it).Example with Caddyβ (Docker imageβ ) (automatic TLS via Let's Encrypt, reverse proxy in the same Docker network):
Caddyfile:
embeddings.example.com {
reverse_proxy embeddings:8000
}
Example with nginx (reverse proxy on the host):
server {
listen 443 ssl;
server_name embeddings.example.com;
ssl_certificate /path/to/cert.pem;
ssl_certificate_key /path/to/key.pem;
location / {
proxy_pass http://127.0.0.1:8000;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_read_timeout 120s;
}
}
Fresh persistent installs auto-generate an EMBED_API_KEY. Display it with docker exec embeddings embed_manage --showkey, or use docker exec embeddings embed_manage --getkey in scripts. For existing installs without a key, set EMBED_API_KEY in your env file to enable authentication.
To update the Docker image and container, first downloadβ the latest version:
docker pull hwdsl2/embeddings-server
If the Docker image is already up to date, you should see:
Status: Image is up to date for hwdsl2/embeddings-server:latest
Otherwise, it will download the latest version. Remove and re-create the container:
docker rm -f embeddings
# Then re-run the docker run command from Quick start with the same volume and port.
Your downloaded models are preserved in the embeddings-data volume.
Embeddings can be used as the embedding service in a broader self-hosted AI setup.
For full and lightweight Docker Compose stacks, manual docker run examples, and voice/RAG/MCP pipeline examples with SpeakCrate, EmbedCrate, GatewayCrate, InferCrate, ParseCrate, and UplinkCrate, see Self-Hosted AI Stackβ .
See Usage countsβ .
ghcr.io/huggingface/text-embeddings-inference:cpu-latest (Debian)/v1/embeddings endpoint (served directly by TEI; supported fields depend on TEI)/rerank endpoint via a second process loaded with a cross-encoder model/var/lib/embeddings (Docker volume)huggingface_hub) for pre-download via embed_manage --pullmodelNote: The software components inside the pre-built image (such as Hugging Face TEI and its dependencies) are under the respective licenses chosen by their respective copyright holders. As for any pre-built image usage, it is the image user's responsibility to ensure that any use of this image complies with any relevant licenses for all software contained within.
Copyright (C) 2026 Lin Song
This work is licensed under the MIT Licenseβ .
Hugging Face Text Embeddings Inference (TEI) is Copyright (C) Hugging Face, Inc., and is distributed under the Apache License 2.0β .
This project is an independent Docker setup for Hugging Face TEI and is not affiliated with, endorsed by, or sponsored by Hugging Face, Inc.
Content type
Image
Digest
sha256:1ba12d253β¦
Size
266.1 MB
Last updated
2 days ago
docker pull hwdsl2/embeddings-server