Skip to content

Make backend cpu only as GPUs are used strictly by other containers, … - #33

Open
ishaanbhela-ai wants to merge 1 commit into
lending_poc/mainfrom
fix/DockerBuildComposeAndSuryaOCR
Open

ishaanbhela-ai wants to merge 1 commit into
lending_poc/mainfrom
fix/DockerBuildComposeAndSuryaOCR

Conversation

@ishaanbhela-ai

@ishaanbhela-ai ishaanbhela-ai commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Optimizes the Docker build workflow and resolves initialization/compilation issues for the GPU stack.

Key Changes

  • Surya Inference Optimization: Switched from building llama.cpp from source (~25 min build time) to the official pre-built ghcr.io/ggml-org/llama.cpp:server-cuda image (builds in seconds).
  • Hugging Face Download Fix: Added User-Agent headers and minimum file-size validation in entrypoint.sh to prevent Hugging Face CDN 403 blocks and corrupt GGUF downloads.
  • Backend Service Simplification: Removed redundant backend-gpu profile. Unified backend to use lightweight CPU PyTorch (GPU: "0") since inference is offloaded to Surya/Ollama, cutting image build time to <1 min.
  • Linux & Cross-Platform Compatibility: Standardized all storage on native Docker named volumes (ollama_models, surya_models, hf_cache, pgdata) without host OS drive-path dependencies.

Testing

  • Verified docker compose --profile gpu build builds cleanly in seconds.
  • Verified all 5 containers start and reach healthy states.
  • Verified Surya OCR inference loads onto GPU (CUDA) and answers /ocr/health.
  • Tested document extraction flow (/extract).

…Remove python from surya container as it is not needed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Partial model downloads can be accepted and launched as corrupt GGUF files.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 Medium severity

Open (1)
What changed in this PR

Updates the Docker OCR stack to use prebuilt llama.cpp images and a CPU-only backend while GPU workloads run in dedicated services.

Changes:

  • Adds Hugging Face headers and model-size validation.
  • Replaces source-built llama.cpp with official images.
  • Simplifies backend services and uses native Docker volumes.
File Summary
lending-poc/​surya-inference/​entrypoint.sh Model download and launch logic. Moderate issue (3 votes): interrupted downloads can leave a partial GGUF at the final path; download to a temporary path, validate it, then atomically rename it.
lending-poc/​surya-inference/​Dockerfile Uses prebuilt llama.cpp images.
lending-poc/​docker-compose.yml Configures the CPU backend, GPU inference services, and named volumes.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +15 to +18
if [ ! -f "$MODEL_PATH" ] || [ $(wc -c < "$MODEL_PATH" 2>/dev/null || echo 0) -lt 10000000 ]; then
rm -f "$MODEL_PATH"
echo "Downloading ${SURYA_GGUF_MODEL_FILE} from ${SURYA_GGUF_REPO}..."
curl -fL -o "$MODEL_PATH" "https://huggingface.co/${SURYA_GGUF_REPO}/resolve/main/${SURYA_GGUF_MODEL_FILE}"
curl -fL -H "User-Agent: Mozilla/5.0" -o "$MODEL_PATH" "https://huggingface.co/${SURYA_GGUF_REPO}/resolve/main/${SURYA_GGUF_MODEL_FILE}"
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants