Custom Service
Run ML models and microservices as Docker containers and expose them as tools in DeepFellow.
The Custom Service lets you run any containerized application — an ML model, a microservice, or any HTTP API — directly on DeepFellow Infra. Once installed, the service is accessible through the /custom/{prefix}/{path} endpoint and can be connected to an LLM as a tool.
Available Models
The following models are ready to install from the Custom Service panel in DeepFellow Infra Web Panel:
| Model ID | Description |
|---|---|
lemmatizer | Multilingual text lemmatization API. |
bentoml/example-summarization | Text summarization service built with BentoML. |
deepfellow-bge-m3 | High-performance multilingual text embeddings model for sparse vectors. |
deepfellow-finetune | LLM LoRA fine-tuning framework with an OpenAI-compatible fine-tuning job queue. |
doc_chunker | Converts many file types (PDF, XLSX, MP3, and others) into text chunks suitable for vector search. |
easyOCR | Optical character recognition service that extracts text from images and binary data. |
To install one of these models:
Add Your Own Model
You can run any containerized ML model or microservice in the Custom Service. This section walks through a complete example using an Iris species classifier.
Requirements
- DeepFellow Infra installed and running
- Docker installed
- uv
1. Train the Model
Create and save a trained ML model as a .pkl file. The example uses scikit-learn:
from sklearn.datasets import load_iris
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
import joblib
iris = load_iris()
X, y = iris.data, iris.target
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = RandomForestClassifier(n_estimators=100, random_state=42)
model.fit(X_train, y_train)
joblib.dump(model, "iris_model.pkl")Run it with:
uv run python create_iris_model.py2. Create a Microservice
Wrap the model in a web API. The example uses FastAPI on port 8001:
from fastapi import FastAPI
import joblib
from pydantic import BaseModel
from typing import List
import numpy
app = FastAPI()
model = joblib.load("iris_model.pkl")
class InputData(BaseModel):
features: List[float]
@app.post("/predict")
def predict(data: InputData) -> dict[str, int]:
prediction: numpy.ndarray = model.predict([data.features])
return {"prediction": prediction.tolist()[0]}3. Write a Dockerfile
Package the microservice in a Docker image:
FROM python:3.11-slim
COPY --from=ghcr.io/astral-sh/uv:latest /uv /uvx /bin/
WORKDIR /app
COPY pyproject.toml uv.lock ./
RUN uv sync --frozen --no-dev
COPY . .
EXPOSE 8001
CMD ["uv", "run", "uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8001"]4. Build the Docker Image
docker build -t iris-model .You are not limited to locally built images. In the Docker image field you can also enter an image tag from Docker Hub or GitHub Container Registry.
5. Install on DeepFellow Infra
iris-model, Docker image iris-model, Docker image port 8001.iris-model as the endpoint prefix when prompted.6. Verify
Send a request to the DeepFellow Infra endpoint:
curl -X POST 'http://DEEPFELLOW_INFRA_HOST/custom/iris-model/predict' \
-H 'Authorization: Bearer DEEPFELLOW_INFRA_API_KEY' \
-H 'Content-Type: application/json' \
-d '{"features": [5.1, 3.5, 1.4, 0.2]}'Expected response:
{"prediction": 0}Use as a Tool
To expose this model to an LLM, add it as a custom-mcp tool. See Using Tools for the full schema and a complete example using this iris classifier.
Edit and Duplicate
Open the ⋮ menu on an installed model's card and click Edit Settings to change its configuration: the endpoint prefix, the Docker image, environment variables, or any other install-time option. DeepFellow uninstalls the running model, applies the new settings, and reinstalls it automatically.
- The Model ID cannot be changed from the Edit dialog. To register the model under a different ID, use Duplicate instead.
- If the endpoint prefix you enter is already used by another installed model, DeepFellow rejects the edit before uninstalling anything.
- Saving an edit with no changes closes the dialog without uninstalling and reinstalling the model.
To create an independent copy of a model instead of changing the original, select Save as a new service in the Edit dialog before saving. DeepFellow installs the copy immediately, under the ID and prefix you provide, with the same settings as the original.
Proxy Timeout
Every custom-service install form, whether for a catalog model or a model you add yourself, includes a Seconds the proxy waits for a response from this service before timing out field. It controls how long the proxy waits for that specific service to respond before timing out the request.
Leave the field empty to use the standard_proxy_timeout_seconds setting (600 seconds by default). Set it to a higher value for a service with a long processing time, such as doc_chunker converting a large file, or to a lower value for a fast service, so a failure is detected sooner. The value you enter applies only to that one installation and doesn't change the global default used by other services.
We use cookies on our website. We use them to ensure proper functioning of the site and, if you agree, for purposes such as analytics, marketing, and targeting ads.