ADR-056: Chat API Client Cancellation and Timeout Handling

Status: Partially Implemented (Phase 0 complete) Date: 2025-12-12 (Proposed), 2025-12-14 (Phase 0 Implemented) Context: Fixing Issue #2 - Client cancellation doesn't stop server processing

Problem Statement

When a client disconnects (Ctrl+C) from /v1/chat/completions, two critical issues occur:

Issue 1: Client Request Hangs

Symptom: Client request never completes, hangs indefinitely Impact: Poor UX, resource leaks on client side

Issue 2: Server Keeps Processing

Symptom: LangChain4j LLM call continues running after client disconnect Impact: Wasted compute, wasted API costs, zombie threads

Error During Shutdown

ERROR [io.qua.ver.htt.run.QuarkusErrorHandler] HTTP Request to /v1/chat/completions failed
[Error Occurred After Shutdown]: java.lang.NullPointerException:
Cannot invoke "io.quarkus.arc.ArcContainer.getActiveContext..." because "<local2>" is null

Root Cause Analysis

1. Non-Streaming Requests (Blocking Call)

Location: ChatApiResource.java:91

ChatResponse response = agentService.chat(request, agentContext);

Problems:

2. LangChain4j Integration (No Cancellation)

Location: ChatAgentService.java:282

String agentResponse = chatAgent.chat(userMessage);

Problems:

3. Streaming Requests (Output Failure)

Location: ChatApiResource.java:169-208

StreamingOutput stream = output -> {
    try (var writer = new java.io.OutputStreamWriter(output)) {
        agentService.chatStream(request, agentContext, new StreamCallback() {
            @Override
            public void onChunk(StreamChunk chunk) {
                writer.write("data: " + json + "\n\n");  // ⚠️ Fails after disconnect
                writer.flush();
            }
        });
    }
};

Problems:

4. SmallRye Context Manager NPE (CRITICAL - Newly Discovered)

Location: ChatAgentService.java:289 → JaxRsHttpClient.execute() (LangChain4j)

Root Cause Stack:

java.lang.NullPointerException: Cannot invoke
"io.smallrye.context.SmallRyeContextManager.defaultThreadContext()"
because the return value of "io.smallrye.context.SmallRyeContextManagerProvider.getManager()" is null
    at SmallRyeThreadContext.getCurrentThreadContextOrDefaultContexts
    at DefaultContextPropagationInterceptor.getThreadContext
    at Infrastructure.decorate
    at UniOnItem.transform
    at ClientSendRequestHandler.createRequest
    at JaxRsHttpClient.execute (LangChain4j)
    at OllamaClient.chat
    at ChatAgent.chat
    at ChatAgentService.executeChat:289

Problem:

Impact: This is the actual production NPE - occurs before the Arc container NPE

5. Shutdown Exception Mapper Failure (SECONDARY)

Location: ExceptionMappers.java (during shutdown)

Problem:

Impact: This NPE only occurs if SmallRye NPE (#4) gets through

Decision

Implement multi-layer cancellation and timeout handling:

1. Add Request Timeout (All Requests)

2. Detect Client Disconnect (Streaming Only)

3. Add Graceful Shutdown Handling

4. Add Interruption Support (Best Effort)

Implementation

Phase 0: Prevent LLM Calls During Shutdown (IMPLEMENTED - v1.66.0)

File: src/main/java/org/idempiere/cli/chatapi/agent/ChatAgentService.java

Problem: SmallRye Context Manager NPE occurs INSIDE the LLM call before our catch block runs.

Solution: Check for shutdown BEFORE calling chatAgent.chat() or chatAgent.streamChat().

private ChatResponse executeChat(ChatRequest request, AgentContext context,
                                 ModelRouter.RoutingResult routing, String requestId) {
    String model = request.model() != null ? request.model() : defaultModel;

    try {
        // Check if shutting down BEFORE making LLM call
        // This prevents SmallRye Context Manager NPE during shutdown
        if (isShuttingDown()) {
            LOG.debugf("Request %s rejected - service is shutting down", requestId);
            throw new RuntimeException("Service is shutting down");
        }

        // Extract the last user message to send to the agent
        ChatMessage lastMessage = request.messages().get(request.messages().size() - 1);
        String userMessage = lastMessage.content();

        // Call the LangChain4j ChatAgent (SAFE - shutdown checked above)
        String agentResponse = chatAgent.chat(userMessage);

        // ... rest of method
    } catch (Exception e) {
        // Still check in catch block for errors DURING execution
        if (isShuttingDown()) {
            LOG.debugf("Agent execution interrupted during shutdown for request %s", requestId);
            throw new RuntimeException("Service is shutting down", e);
        }
        // ... normal error handling
    }
}

private void executeStreamingChat(ChatRequest request, AgentContext context,
                                  ModelRouter.RoutingResult routing, String requestId,
                                  StreamCallback callback) {
    try {
        // Check if shutting down BEFORE making LLM streaming call
        if (isShuttingDown()) {
            LOG.debugf("Streaming request %s rejected - service is shutting down", requestId);
            callback.onError(new RuntimeException("Service is shutting down"));
            return;
        }

        // Safe to call LangChain4j streaming (shutdown checked above)
        chatAgent.streamChat(userMessage).subscribe().with(...);
    } catch (Exception e) {
        // ... error handling
    }
}

/**
 * Enhanced shutdown detection - checks BOTH Arc and SmallRye managers.
 */
private boolean isShuttingDown() {
    try {
        // Check Arc CDI container
        var container = Arc.container();
        if (container == null) {
            return true;
        }

        // Check SmallRye Context Manager (used by LangChain4j JAX-RS client)
        var contextManager = io.smallrye.context.SmallRyeContextManagerProvider.getManager();
        if (contextManager == null) {
            return true;
        }

        return false;
    } catch (Exception e) {
        // If we can't access containers, assume we're shutting down
        return true;
    }
}

Result:

Phase 1: Add Timeout Wrapper (PROPOSED)

File: src/main/java/org/idempiere/cli/chatapi/agent/ChatAgentService.java

@ConfigProperty(name = "chat-api.agent.request-timeout-seconds", defaultValue = "120")
int requestTimeoutSeconds;

private ChatResponse executeChat(ChatRequest request, AgentContext context,
                                 ModelRouter.RoutingResult routing, String requestId) {
    // Wrap in CompletableFuture with timeout
    CompletableFuture<String> future = CompletableFuture.supplyAsync(() -> {
        try {
            ChatMessage lastMessage = request.messages().get(request.messages().size() - 1);
            return chatAgent.chat(lastMessage.content());
        } catch (Exception e) {
            throw new CompletionException("Agent execution failed", e);
        }
    });

    try {
        // Wait with timeout
        String agentResponse = future.get(requestTimeoutSeconds, TimeUnit.SECONDS);

        // Build response (existing code)
        // ...

    } catch (TimeoutException e) {
        future.cancel(true);  // Interrupt the thread
        LOG.warnf("Request %s timed out after %d seconds", requestId, requestTimeoutSeconds);
        throw new RuntimeException("Request timed out after " + requestTimeoutSeconds + " seconds");
    } catch (InterruptedException e) {
        Thread.currentThread().interrupt();
        throw new RuntimeException("Request interrupted", e);
    } catch (ExecutionException e) {
        throw new RuntimeException("Agent execution failed", e.getCause());
    }
}

Phase 2: Detect Client Disconnect in Streaming

File: src/main/java/org/idempiere/cli/chatapi/agent/ChatAgentService.java

private void executeStreamingChat(ChatRequest request, AgentContext context,
                                  ModelRouter.RoutingResult routing, String requestId,
                                  StreamCallback callback) {
    String model = request.model() != null ? request.model() : defaultModel;

    try {
        ChatMessage lastMessage = request.messages().get(request.messages().size() - 1);
        String userMessage = lastMessage.content();

        int inputTokens = estimateInputTokens(request);
        final int[] outputTokenCount = {0};

        callback.onChunk(StreamChunk.start(requestId, model, ChatMessage.Role.ASSISTANT));

        // Track subscription for cancellation
        AtomicReference<Subscription> subscriptionRef = new AtomicReference<>();
        AtomicBoolean clientDisconnected = new AtomicBoolean(false);

        chatAgent.streamChat(userMessage)
            .subscribe().with(
                subscription -> subscriptionRef.set(subscription),

                // On each token
                token -> {
                    if (clientDisconnected.get()) {
                        // Client disconnected, stop processing
                        return;
                    }

                    try {
                        callback.onChunk(StreamChunk.content(requestId, model, 0, token));
                        outputTokenCount[0] += token.length() / 4;
                    } catch (Exception e) {
                        // Client disconnect detected (IOException from writer)
                        LOG.debugf("Client disconnected for request %s, cancelling stream", requestId);
                        clientDisconnected.set(true);

                        // Cancel subscription to stop LangChain4j processing
                        Subscription sub = subscriptionRef.get();
                        if (sub != null) {
                            sub.cancel();
                        }
                    }
                },

                // On error
                error -> {
                    if (!clientDisconnected.get()) {
                        LOG.errorf(error, "Streaming chat failed for request %s", requestId);
                        callback.onError(error);
                    }
                },

                // On complete
                () -> {
                    if (clientDisconnected.get()) {
                        LOG.debugf("Stream completed after client disconnect for request %s", requestId);
                        return;
                    }

                    TokenUsage usage = TokenUsage.of(
                        inputTokens,
                        outputTokenCount[0],
                        model,
                        costCalculator.calculateCost(model, inputTokens, outputTokenCount[0], 0)
                    );

                    callback.onChunk(StreamChunk.finish(requestId, model, "stop", usage));
                    callback.onComplete(usage);

                    costGuard.recordSpending(
                        usage.estimatedCostUsd(),
                        context.getClientId(),
                        context.getUserId()
                    );
                }
            );

    } catch (Exception e) {
        LOG.errorf(e, "Failed to start streaming chat for request %s", requestId);
        callback.onError(e);
    }
}

Phase 3: Improve Shutdown Handling

File: src/main/java/org/idempiere/cli/chatapi/api/ExceptionMappers.java

@Provider
public class GenericExceptionMapper implements ExceptionMapper<Exception> {

    @Override
    public Response toResponse(Exception exception) {
        // Check if shutting down BEFORE accessing CDI
        if (isShuttingDown()) {
            return Response.status(Response.Status.SERVICE_UNAVAILABLE)
                .entity(ChatResponse.error(
                    "Service is shutting down",
                    "shutdown_error",
                    null
                ))
                .build();
        }

        // Normal exception handling (can access CDI beans safely)
        // ...
    }

    private boolean isShuttingDown() {
        try {
            var container = Arc.container();
            return container == null;
        } catch (Exception e) {
            return true;
        }
    }
}

Phase 4: Add Configuration

File: src/main/resources/application.properties

# Chat API Agent Configuration
chat-api.agent.request-timeout-seconds=120
chat-api.agent.shutdown-grace-period-seconds=30

Trade-offs

Timeout Approach

Pros:

Cons:

Decision: Use 120 seconds default (2 minutes), make configurable

Client Disconnect Detection

Pros:

Cons:

Decision: Implement best-effort cancellation, document limitations

Thread Interruption

Pros:

Cons:

Decision: Use future.cancel(true) but document it's best-effort

Configuration

Property Default Description
chat-api.agent.request-timeout-seconds 120 Maximum time for LLM request
chat-api.agent.shutdown-grace-period-seconds 30 Time to wait for in-flight requests during shutdown

Testing

Test 1: Non-Streaming Timeout

# Send request that takes > 120 seconds
curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"claude-sonnet-4","messages":[{"role":"user","content":"..."}]}'

# Expected: 408 Request Timeout after 120 seconds

Test 2: Streaming Client Disconnect

# Start streaming request, Ctrl+C after 5 seconds
curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"claude-sonnet-4","messages":[...],"stream":true}'

# Expected: Server logs "Client disconnected", cancels processing

Test 3: Graceful Shutdown

# Start long request, then Ctrl+C server
curl -X POST http://localhost:8080/v1/chat/completions &
sleep 2
pkill -INT java  # Send SIGINT to server

# Expected: Server waits up to 30 seconds for request to complete

Limitations

  1. LangChain4j API Calls Not Interruptible

    • Anthropic/OpenAI HTTP calls may continue even after timeout/cancellation
    • Best effort: we stop processing locally but API call may complete
    • API costs still incurred for interrupted requests
  2. Reactive Streams Cancellation

    • LangChain4j may not support cancellation on all providers
    • Subscription.cancel() is advisory, not guaranteed
  3. Partial Results Lost

    • On timeout/disconnect, partial LLM responses are discarded
    • No mechanism to resume or recover partial progress

Success Criteria

Phase 0 (IMPLEMENTED - v1.66.0)

Phase 1-4 (PROPOSED - Future)

Monitoring

Add metrics to track:

References


Version: v1.66.0 (Phase 0 implemented), v1.67.0+ (Phase 1-4 proposed) Status: Partially Implemented

Implementation History

Next Steps

Phase 1-4 (timeout wrapper, client disconnect detection, etc.) remain as proposed enhancements for future releases.

Path: /docs/developers/architecture/idempiere-hub/056-chat-api-cancellation-and-timeout