Make backend cpu only as GPUs are used strictly by other containers, … - #33
Open
ishaanbhela-ai wants to merge 1 commit into
Open
ishaanbhela-ai wants to merge 1 commit into
ishaanbhela-ai wants to merge 1 commit into
Conversation
…Remove python from surya container as it is not needed
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Partial model downloads can be accepted and launched as corrupt GGUF files.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 1
What changed in this PR
Updates the Docker OCR stack to use prebuilt llama.cpp images and a CPU-only backend while GPU workloads run in dedicated services.
Changes:
- Adds Hugging Face headers and model-size validation.
- Replaces source-built llama.cpp with official images.
- Simplifies backend services and uses native Docker volumes.
| File | Summary |
|---|---|
lending-poc/surya-inference/entrypoint.sh |
Model download and launch logic. Moderate issue (3 votes): interrupted downloads can leave a partial GGUF at the final path; download to a temporary path, validate it, then atomically rename it. |
lending-poc/surya-inference/Dockerfile |
Uses prebuilt llama.cpp images. |
lending-poc/docker-compose.yml |
Configures the CPU backend, GPU inference services, and named volumes. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+15
to
+18
| if [ ! -f "$MODEL_PATH" ] || [ $(wc -c < "$MODEL_PATH" 2>/dev/null || echo 0) -lt 10000000 ]; then | ||
| rm -f "$MODEL_PATH" | ||
| echo "Downloading ${SURYA_GGUF_MODEL_FILE} from ${SURYA_GGUF_REPO}..." | ||
| curl -fL -o "$MODEL_PATH" "https://huggingface.co/${SURYA_GGUF_REPO}/resolve/main/${SURYA_GGUF_MODEL_FILE}" | ||
| curl -fL -H "User-Agent: Mozilla/5.0" -o "$MODEL_PATH" "https://huggingface.co/${SURYA_GGUF_REPO}/resolve/main/${SURYA_GGUF_MODEL_FILE}" |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Summary
Optimizes the Docker build workflow and resolves initialization/compilation issues for the GPU stack.
Key Changes
llama.cppfrom source (~25 min build time) to the official pre-builtghcr.io/ggml-org/llama.cpp:server-cudaimage (builds in seconds).User-Agentheaders and minimum file-size validation inentrypoint.shto prevent Hugging Face CDN 403 blocks and corrupt GGUF downloads.backend-gpuprofile. Unifiedbackendto use lightweight CPU PyTorch (GPU: "0") since inference is offloaded to Surya/Ollama, cutting image build time to <1 min.ollama_models,surya_models,hf_cache,pgdata) without host OS drive-path dependencies.Testing
docker compose --profile gpu buildbuilds cleanly in seconds.CUDA) and answers/ocr/health./extract).