-
Notifications
You must be signed in to change notification settings - Fork 10
Expand file tree
/
Copy path.env.example
More file actions
270 lines (245 loc) · 10.8 KB
/
Copy path.env.example
File metadata and controls
270 lines (245 loc) · 10.8 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
# ============================================================================
# CODEBASE INDEX CLI - GLOBAL CONFIGURATION
# ============================================================================
# Copy this file to ./.env (in the project root) and edit it with your
# credentials. This file is used as global configuration for all workspaces.
#
# IMPORTANT: You only need to configure ONE embeddings section
# (OpenAI, OpenAI-compatible, or Ollama) depending on which provider you use.
# ============================================================================
# ----------------------------------------------------------------------------
# GENERAL SETTINGS
# ----------------------------------------------------------------------------
# Logging level: debug, info, warn, error
LOG_LEVEL=info
# Workspace path (optional, defaults to current directory)
# CODEBASE_WORKSPACE_PATH=/abs/path/to/workspace
# JSON config file (optional, if you prefer JSON over ENV)
# CODEBASE_INDEX_CONFIG=./codebase-index.config.json
# ============================================================================
# VECTOR STORE CONFIGURATION
# ============================================================================
# The vector store is determined by which command you use:
# - codebase → Uses Qdrant (remote server)
# - codesql → Uses SQLite-vec (local database)
#
# Configure BOTH options below. The appropriate one will be used based on
# which command you run. This allows you to use different vector stores
# for different projects without changing configuration.
# ============================================================================
# ----------------------------------------------------------------------------
# SQLite-vec Configuration (used by 'codesql' command)
# ----------------------------------------------------------------------------
# Local database stored in your project
# ✅ No external services required
# ✅ Portable, can be committed with your project
# ✅ Best for small/medium projects
#
# Usage: codesql -start .
SQLITE_DB_PATH=.codebase/vectors.db
SQLITE_SEARCH_MIN_SCORE=0.4
SQLITE_SEARCH_MAX_RESULTS=50
# ----------------------------------------------------------------------------
# Qdrant Configuration (used by 'codebase' command)
# ----------------------------------------------------------------------------
# Requires Qdrant server running:
# docker run -p 6333:6333 qdrant/qdrant
# ✅ High performance
# ✅ Scalable to millions of vectors
# ✅ Best for large projects or production
#
# Usage: codebase -start .
QDRANT_URL=http://localhost:6333
QDRANT_API_KEY=
QDRANT_COLLECTION=
QDRANT_SEARCH_MIN_SCORE=0.4
QDRANT_SEARCH_MAX_RESULTS=50
# ============================================================================
# EMBEDDINGS PROVIDER SELECTION
# ============================================================================
# Choose which embeddings provider to use: "openai", "openai-compatible", or "ollama"
#
# HOW IT WORKS:
# - Set EMBED_PROVIDER=openai → Uses official OpenAI API
# - Set EMBED_PROVIDER=openai-compatible → Uses custom OpenAI-compatible API (Nebius, Together, etc.)
# - Set EMBED_PROVIDER=ollama → Uses local Ollama server
#
# You can configure ALL options below, but only the one specified in
# EMBED_PROVIDER will be used. This allows you to switch between them easily.
#
# NEW FEATURE: Vector-Store-Specific Embedders
# ============================================
# You can now use DIFFERENT embedders for SQLite vs Qdrant!
#
# Use prefixes to specify embedder per vector store:
# - SQLITE_EMBED_* → Used only when vector store is SQLite
# - QDRANT_EMBED_* → Used only when vector store is Qdrant
# - EMBED_* (no prefix) → Global fallback for both
#
# Example: Use small/fast model for SQLite, large/accurate for Qdrant:
# SQLITE_EMBED_MODEL=text-embedding-3-small
# QDRANT_EMBED_MODEL=Qwen/Qwen3-Embedding-8B
#
# If vector-store-specific variables are not set, falls back to global EMBED_* variables.
#
# Default: openai-compatible (for custom APIs like Nebius)
EMBED_PROVIDER=openai-compatible
# ----------------------------------------------------------------------------
# OPTION 1: OpenAI (Official)
# ----------------------------------------------------------------------------
# Uses official OpenAI API
# Requires: OPENAI_API_KEY
# Available models:
# - text-embedding-3-small (1536 dims, fast, economical, 8192 token limit)
# - text-embedding-3-large (3072 dims, better quality, 8192 token limit)
# - text-embedding-ada-002 (1536 dims, legacy)
OPENAI_API_KEY=sk-your-openai-key-here
OPENAI_EMBED_MODEL=text-embedding-3-small
OPENAI_MAX_BATCH=60
OPENAI_MAX_TOKENS=8192
# ----------------------------------------------------------------------------
# OPTION 2: OpenAI-Compatible (Recommended for custom APIs)
# ----------------------------------------------------------------------------
# Uses any OpenAI-compatible API (Nebius, Together, etc.)
# Requires: EMBED_BASE_URL, EMBED_API_KEY, EMBED_MODEL
#
# Example - Nebius AI Studio:
# - Base URL: https://api.studio.nebius.com/v1/
# - Model: Qwen/Qwen3-Embedding-8B (4096 dims)
# - API Key: your Nebius key
EMBED_BASE_URL=https://api.studio.nebius.com/v1/
EMBED_API_KEY=your-api-key-here
EMBED_MODEL=Qwen/Qwen3-Embedding-8B
EMBED_MAX_BATCH=60
EMBED_MAX_TOKENS=8192
# ----------------------------------------------------------------------------
# OPTION 3: Ollama (Local, Self-hosted)
# ----------------------------------------------------------------------------
# Uses Ollama running locally
# Requires: Ollama installed and running (ollama serve)
# Available models: nomic-embed-text, mxbai-embed-large, etc.
# IMPORTANT: You must specify the model dimension
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_MODEL=nomic-embed-text
OLLAMA_EMBED_DIMENSION=768
OLLAMA_MAX_TOKENS=8192
# ----------------------------------------------------------------------------
# BATCH CONFIGURATION (Advanced)
# ----------------------------------------------------------------------------
# EMBED_MAX_BATCH: Maximum number of texts to send in one API request
# - Higher = faster indexing but more memory usage
# - Lower = slower but safer for rate limits
# - Recommended: 50-100 for most models
#
# EMBED_MAX_TOKENS: Maximum total tokens per API request
# - CRITICAL: Must match your model's token limit!
# - text-embedding-3-small: 8192 tokens
# - text-embedding-3-large: 8192 tokens
# - Qwen/Qwen3-Embedding-8B: 32768 tokens (check your provider)
# - If you get "maximum context length" errors, REDUCE this value
# - Default: 8192 (safe for most models)
#
# Example error that means you need to reduce EMBED_MAX_TOKENS:
# "This model's maximum context length is 8192 tokens, however you requested 16990 tokens"
#
# Solution: Set EMBED_MAX_TOKENS to a value BELOW the model's limit (e.g., 7000 for safety margin)
# ----------------------------------------------------------------------------
# VECTOR-STORE-SPECIFIC EMBEDDERS (Optional)
# ----------------------------------------------------------------------------
# Use different embedders for SQLite vs Qdrant
# These override the global EMBED_* variables for their respective vector store
#
# Example 1: Different models for each store
# SQLITE_EMBED_PROVIDER=openai
# SQLITE_EMBED_MODEL=text-embedding-3-small
# SQLITE_EMBED_API_KEY=sk-...
# SQLITE_EMBED_DIMENSION=1536
#
# QDRANT_EMBED_PROVIDER=openai-compatible
# QDRANT_EMBED_MODEL=Qwen/Qwen3-Embedding-8B
# QDRANT_EMBED_API_KEY=nebius-key
# QDRANT_EMBED_BASE_URL=https://api.studio.nebius.com/v1/
# QDRANT_EMBED_DIMENSION=4096
#
# Example 2: Use Ollama for SQLite (local/fast), OpenAI for Qdrant (cloud/accurate)
# SQLITE_EMBED_PROVIDER=ollama
# SQLITE_OLLAMA_MODEL=nomic-embed-text
# SQLITE_OLLAMA_BASE_URL=http://localhost:11434
# SQLITE_OLLAMA_EMBED_DIMENSION=768
#
# QDRANT_EMBED_PROVIDER=openai
# QDRANT_OPENAI_API_KEY=sk-...
# QDRANT_OPENAI_EMBED_MODEL=text-embedding-3-large
# QDRANT_OPENAI_EMBED_DIMENSION=3072
# ============================================================================
# ADVANCED OPTIONS
# ============================================================================
# ----------------------------------------------------------------------------
# Indexing
# ----------------------------------------------------------------------------
# Cache file path (file hashes)
INDEXER_CACHE_PATH=.codebase/cache.json
# Batch size for embeddings (adjust based on API limits)
INDEXER_BATCH_SIZE=64
# Maximum file size to index (in bytes, 1MB = 1048576)
INDEXER_MAX_FILE_SIZE_BYTES=1048576
# File patterns to index (glob patterns, comma-separated)
# Example: src/**/*.ts,src/**/*.tsx,lib/**/*.js
INDEXER_FILE_GLOBS=src/**/*.ts,src/**/*.tsx
# ----------------------------------------------------------------------------
# File Watcher
# ----------------------------------------------------------------------------
# Enable file change watcher
INDEXER_WATCH_ENABLED=true
# Wait time before processing changes (in milliseconds)
INDEXER_WATCH_DEBOUNCE_MS=500
# ----------------------------------------------------------------------------
# Tree-sitter (Semantic Parsing)
# ----------------------------------------------------------------------------
# Enable semantic parsing with tree-sitter
# - true: Uses AST to split code by functions/classes
# - false: Uses simple regex (default, faster)
# Supports 29+ languages (TypeScript, Python, Rust, etc.)
USE_TREE_SITTER=false
# ----------------------------------------------------------------------------
# Git Commit Tracking (Experimental)
# ----------------------------------------------------------------------------
# Track git commits in the workspace
# - true: Monitors .git directory for new commits
# - false: Git tracking disabled (default)
#
# When enabled, the watcher will detect:
# - New commits (local commits)
# - Pulled commits (git pull)
# - Branch changes (git checkout)
# - Merges (git merge)
#
# When a commit is detected, the system will:
# 1. Extract commit metadata (author, date, message, diff)
# 2. Generate a structured prompt with codebase context
# 3. Send prompt to LLM for semantic analysis (if configured)
# 4. Index the LLM response in the vector store
# 5. Enable semantic search of commit history
#
TRACK_GIT=false
# ----------------------------------------------------------------------------
# Git Commit LLM Analysis (Requires TRACK_GIT=true)
# ----------------------------------------------------------------------------
# Configure LLM provider for analyzing git commits
# Supported providers: openai-compatible
#
# When configured, each commit will be:
# - Analyzed by the LLM for semantic understanding
# - Indexed in the vector store for semantic search
# - Searchable alongside your codebase
#
# Example configuration:
# TRACK_GIT_LLM_PROVIDER=openai-compatible
# TRACK_GIT_LLM_ENDPOINT=http://localhost:4141/v1
# TRACK_GIT_LLM_MODEL=gpt-4.1
# TRACK_GIT_LLM_API_KEY=sk-your-key-here
#
TRACK_GIT_LLM_PROVIDER=
TRACK_GIT_LLM_ENDPOINT=
TRACK_GIT_LLM_MODEL=
TRACK_GIT_LLM_API_KEY=