An OpenAI-compatible API server for running RKLLM (Rockchip Large Language Model) models on RK3588 platform.
A modular and well-structured server to run inference on RKLLM models, with proper support for:
- OpenAI-compatible API endpoints
- Tool/Function calling
- Streaming and non-streaming responses
- Conversation history management
- WebUI integration
This project follows a modern Python package structure:
-
rkllm/- Main package directory__init__.py- Package initializationapp.py- Flask application setup and server initializationconfig.py- Centralized configuration moduleapi/- API endpoint implementationsroutes.py- OpenAI-compatible API routes
lib/- Library interface coderkllm_runtime.py- Low-level interface to RKLLM native library
utils/- Utility functions and helperstool_calling.py- Tool/function calling utilitiesconversation.py- Conversation history management
-
main.py- Main entry point to start the server -
tests/- Test scriptssimple_tool_test.py- Basic tool calling testtest_tool_calling.py- Comprehensive tool calling teststest_normal_chat.py- Regular chat functionality tests
-
client.py- Client script for interacting with the server -
reference/- Reference materials and documentation -
assets/- Images and other static resources -
src/- Native libraries folder containinglibrkllmrt.so
-
Download an RKLLM model
Obtain your.rkllmmodel file. -
Place the model file
Put your model (e.g.,Qwen3-0.6B-w8a8-opt1-hybrid1-npu3.rkllm) in a directory, for example:../model/Qwen3-0.6B-w8a8-opt1-hybrid1-npu3.rkllm -
Set up environment variables (optional)
You can configure the server using environment variables. For example:export RKLLM_MODEL_PATH="../model/Qwen3-0.6B-w8a8-opt1-hybrid1-npu3.rkllm" export RKLLM_SERVER_PORT=1306
Alternatively, edit the configuration in
rkllm/config.py. -
Start the server
Run the server using the new entry point:python3 main.py
-
Test the server
Run one of the test scripts to verify functionality:# Test basic tool calling python3 tests/simple_tool_test.py # Test comprehensive tool calling features python3 tests/test_tool_calling.py # Test regular chat functionality python3 tests/test_normal_chat.py
-
Use the API
- The server exposes OpenAI-compatible endpoints at
http://0.0.0.0:1306/v1 - Use with any OpenAI client by setting the
base_urlto this address - Key endpoints:
POST /v1/chat/completions- Chat completions APIGET /v1/models- List available models
- The server exposes OpenAI-compatible endpoints at
Implementation Notes:
- The simplified server (
simple_server.py) provides the stability of the original implementation while fitting into the new modular structure - Tool calling and function execution are fully supported in the simplified bridge implementation
- The fully modularized version is still under development
Notes:
- Make sure
librkllmrt.so(the RKLLM runtime) is present in./src/or update the library path in the configuration. - For custom models, adjust the model path in the configuration.
- The server supports both streaming and non-streaming modes.
- Tool/function calling is fully supported with the Qwen3 model.
- Model Collection
The server supports OpenAI-compatible function/tool calling with Qwen3 models. To use tool calling:
import requests
tools = [
{
"type": "function",
"function": {
"name": "calculate",
"description": "Calculate a mathematical expression",
"parameters": {
"type": "object",
"properties": {
"expression": {
"type": "string",
"description": "The mathematical expression to calculate"
}
},
"required": ["expression"]
}
}
}
]
response = requests.post(
"http://localhost:1306/v1/chat/completions",
json={
"model": "gpt-3.5-turbo", # Model name is ignored by the server
"messages": [
{"role": "user", "content": "Calculate 123 + 456"}
],
"tools": tools,
"stream": False
}
)
print(response.json())