Skip to content

Latest commit

 

History

62 Commits

Folders and files

Repository files navigation

Bespoke Nimble Python API library

PyPI version

The Bespoke Nimble Python library provides convenient access to the Bespoke Labs REST API from any Python 3.8+ application. The library includes type definitions for all request params and response fields, and offers both synchronous and asynchronous clients powered by httpx.

It is generated with Stainless.

Documentation

The REST API documentation can be found on docs.bespokelabs.ai. The full API of this library can be found in api.md.

Installation

# install or upgrade from PyPI
pip install --upgrade bespokelabs-nimble

Package migration

This checkout prepares the next bespokelabs-nimble release. The commands below apply once that release is published; PyPI's existing 0.1.0 is the older standalone SDK. To try this change before release, install this checkout with pip install . in a fresh virtual environment.

From bespokelabs

Replace the dependency bespokelabs with bespokelabs-nimble and update imports:

# Before
from bespokelabs import BespokeLabs, AsyncBespokeLabs

# After
from bespokelabs.nimble import Nimble, AsyncNimble

Types and exceptions also move under bespokelabs.nimble, for example from bespokelabs.nimble.types.nimble import Question and from bespokelabs.nimble import APIError. Use client.system_one(...) for Nimble and client.factcheck(...) for batched fact checking. BespokeLabs / AsyncBespokeLabs remain aliases of Nimble / AsyncNimble. The previous client.nimble.system_one(...) and client.minicheck.factcheck.create(...) calls remain available.

Use a fresh virtual environment when migrating an environment that also has Curator or Sandbox installed. The old bespokelabs distribution shares its root bespokelabs/__init__.py with those packages; uninstalling it can remove that shared file. This SDK now owns only bespokelabs/nimble/ and does not export clients from the shared root.

From standalone bespokelabs-nimble==0.1.0

Import Nimble / AsyncNimble from bespokelabs.nimble. The direct client.system_one(...) method remains available, but authentication and question construction differ from the standalone SDK. Use question dictionaries as in the examples below instead of the standalone Noul, Choice, and Score constructors.

The client requires BESPOKE_API_KEY (or api_key=) and defaults to the hosted Bespoke gateway. Custom gateways use BESPOKE_LABS_BASE_URL (or base_url=) and must expose /v1/nimble/systemone; a standalone model server exposing /v1/systemone is not interchangeable. health(), models(), and limits() from the standalone SDK are not provided by this client. Applications that still need that interface should keep bespokelabs-nimble==0.1.0 pinned until they migrate.

Authentication

Set BESPOKE_API_KEY in your environment:

export BESPOKE_API_KEY="your-api-key"

The SDK reads this variable automatically when you create a client:

from bespokelabs.nimble import Nimble

client = Nimble()

You can also pass a key explicitly with api_key:

from bespokelabs.nimble import Nimble

client = Nimble(api_key="your-api-key")

The same options work with AsyncNimble. An explicit api_key takes precedence over BESPOKE_API_KEY. Keep real keys out of source control.

If you keep credentials in a .env file, install python-dotenv and load the file before creating the client; the SDK does not load .env files itself:

from dotenv import load_dotenv
from bespokelabs.nimble import Nimble

load_dotenv()  # Loads BESPOKE_API_KEY from .env into the environment.
client = Nimble()

Upgrading to 0.4.0

Replace auth_token= with api_key= when constructing a client or calling copy() or with_options(). Replace client.auth_token with client.api_key. The BESPOKE_API_KEY environment variable and HTTP authentication header are unchanged, so clients configured solely through the environment need no changes.

Usage

The full API of this library can be found in api.md.

from bespokelabs import nimble

client = nimble.Nimble()  # Reads BESPOKE_API_KEY from the environment.

result = client.system_one(
    state="Please refund the duplicate payment.",
    questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
)
print(result.nouls["refund"].noul)

Nimble

Use Nimble to answer structured questions about text with Noul, Choice, and Score outputs:

from bespokelabs.nimble import Nimble

with Nimble() as client:
    result = client.system_one(
        state="Please refund the duplicate payment.",
        questions={
            "refund": {"type": "noul", "instructions": "Does the customer request a refund?"},
            "department": {
                "type": "choice",
                "instructions": "Which department should handle this request?",
                "criteria": {"billing": "Payments and refunds", "technical": "Software bugs"},
            },
            "urgency": {
                "type": "score",
                "instructions": "Assess operational urgency.",
                "criteria": ["Service works normally", "Partial outage", "Complete outage"],
            },
        },
    )
    print(result.nouls["refund"].noul)
    print(result.choices["department"].choice)
    print(result.scores["urgency"].score)

The default client sends POST /v1/nimble/systemone to https://api.bespokelabs.ai. For a custom deployment, set base_url or BESPOKE_LABS_BASE_URL to a gateway exposing the same route. This route differs from a standalone Nimble model server's /v1/systemone. A custom gateway without the route returns 404; one without a configured model server returns 503.

Noul returns the probability of true. Choice returns a selected option and the probability of each option. Score returns the expected zero-based rubric index: for three levels, its range is 0–2. Confidence measures distribution concentration, not calibrated correctness. There are at most 64 questions per request and 2–26 candidates per Choice or Score. The default model is nimble-latest; pass model= to select another server-supported model.

Answers are typed and available through result.answers, or through result.nouls, result.choices, and result.scores. Question types are exported from bespokelabs.nimble.types.nimble.

The same resource is available on AsyncNimble:

from bespokelabs.nimble import AsyncNimble

async def check_refund():
    async with AsyncNimble() as client:
        result = await client.system_one(
            state="Please refund the duplicate payment.",
            questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
        )
        return result.nouls["refund"].noul

The standard SDK options work for Nimble, including with_options, retries, per-call timeouts, client.with_raw_response.system_one(...), and client.with_streaming_response.system_one(...). Streaming here controls HTTP body reading; it does not produce incremental model answers. Increase the request timeout for slow responses (for example, timeout=180.0). A longer timeout does not make a server wait when it immediately returns an overload or startup error such as 503 or 529; those responses use the SDK's configured retry policy.

Fact checking

from bespokelabs import nimble

with nimble.Nimble() as client:
    response = client.factcheck(
        context="Paris is the capital of France.",
        claims=["France's capital is Paris.", "France's capital is London."],
        effort="medium",
    )
    for result in response.results:
        print(result.claim, result.support_prob)

factcheck accepts 1–64 non-empty claims, each up to 4,000 characters, against up to 400,000 characters of context. It sends one request to /v1/nimble/factcheck. Results preserve input order and duplicate claims. support_prob is the probability that the context supports the whole claim; it is not a calibrated guarantee of truth. supported means the score is above 0.5.

All three effort values are supported:

  • "low": score with the small fact-checking model.
  • "medium" (default): recheck uncertain claims with a larger model.
  • "high": recheck uncertain claims with additional judging and reasoning models.

split_claims=True is the default. Each sentence is checked separately and the lowest support score is returned for the original claim. Set it to False to check each claim whole. Medium and high may return escalated and per-model scores for each claim. If the requested tier cannot finish because its additional models are starting, the API returns a retryable 503 and does not charge the request.

The response includes model, effort, usage, and request_id. Billable input counts the original context once and each original claim once, using the service's shared tokenizer. Internal splitting, chunking and reasoning do not multiply that count; output_tokens is zero because the response contains scores, not generated text. The selected effort determines the input-token rate.

Use await client.factcheck(...) with AsyncNimble. Both clients accept extra_headers, extra_query, and timeout. The effort selects the model; model is not an argument to this method. The legacy client.minicheck.factcheck.create(claim=..., context=...) is unchanged.

Async usage

Simply import AsyncNimble instead of Nimble and use await with each API call:

import asyncio
from bespokelabs.nimble import AsyncNimble

client = AsyncNimble()  # Reads BESPOKE_API_KEY from the environment.


async def main() -> None:
    result = await client.system_one(
        state="Please refund the duplicate payment.",
        questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
    )
    print(result.nouls["refund"].noul)


asyncio.run(main())

Functionality between the synchronous and asynchronous clients is otherwise identical.

Using types

Nested request parameters are TypedDicts. Responses are Pydantic models which also provide helper methods for things like:

  • Serializing back into JSON, model.to_json()
  • Converting to a dictionary, model.to_dict()

Typed requests and responses provide autocomplete and documentation within your editor. If you would like to see type errors in VS Code to help catch bugs earlier, set python.analysis.typeCheckingMode to basic.

Handling errors

When the library is unable to connect to the API (for example, due to network connection problems or a timeout), a subclass of bespokelabs.nimble.APIConnectionError is raised.

When the API returns a non-success status code (that is, 4xx or 5xx response), a subclass of bespokelabs.nimble.APIStatusError is raised, containing status_code and response properties.

All errors inherit from bespokelabs.nimble.APIError.

import bespokelabs.nimble
from bespokelabs.nimble import Nimble

client = Nimble()

try:
    client.system_one(
        state="Please refund the duplicate payment.",
        questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
    )
except bespokelabs.nimble.APIConnectionError as e:
    print("The server could not be reached")
    print(e.__cause__)  # an underlying Exception, likely raised within httpx.
except bespokelabs.nimble.RateLimitError as e:
    print("A 429 status code was received; we should back off a bit.")
except bespokelabs.nimble.APIStatusError as e:
    print("Another non-200-range status code was received")
    print(e.status_code)
    print(e.response)

Error codes are as followed:

Status Code Error Type
400 BadRequestError
401 AuthenticationError
403 PermissionDeniedError
404 NotFoundError
422 UnprocessableEntityError
429 RateLimitError
>=500 InternalServerError
N/A APIConnectionError

Retries

Certain errors are automatically retried 2 times by default, with a short exponential backoff. Connection errors (for example, due to a network connectivity problem), 408 Request Timeout, 409 Conflict, 429 Rate Limit, and >=500 Internal errors are all retried by default.

You can use the max_retries option to configure or disable retry settings:

from bespokelabs.nimble import Nimble

# Configure the default for all requests:
client = Nimble(
    # default is 2
    max_retries=0,
)

# Or, configure per-request:
client.with_options(max_retries=5).nimble.system_one(
    state="Please refund the duplicate payment.",
    questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
)

Timeouts

By default requests time out after 1 minute. You can configure this with a timeout option, which accepts a float or an httpx.Timeout object:

import httpx
from bespokelabs.nimble import Nimble

# Configure the default for all requests:
client = Nimble(
    # 20 seconds (default is 1 minute)
    timeout=20.0,
)

# More granular control:
client = Nimble(
    timeout=httpx.Timeout(60.0, read=5.0, write=10.0, connect=2.0),
)

# Override per-request:
client.with_options(timeout=5.0).nimble.system_one(
    state="Please refund the duplicate payment.",
    questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
)

On timeout, an APITimeoutError is thrown.

Note that requests that time out are retried twice by default.

Advanced

Logging

We use the standard library logging module.

You can enable logging by setting the environment variable BESPOKE_LABS_LOG to info.

$ export BESPOKE_LABS_LOG=info

Or to debug for more verbose logging.

How to tell whether None means null or missing

In an API response, a field may be explicitly null, or missing entirely; in either case, its value is None in this library. You can differentiate the two cases with .model_fields_set:

if response.my_field is None:
  if 'my_field' not in response.model_fields_set:
    print('Got json like {}, without a "my_field" key present at all.')
  else:
    print('Got json like {"my_field": null}.')

Accessing raw response data (e.g. headers)

The "raw" Response object can be accessed by prefixing .with_raw_response. to any HTTP method call, e.g.,

from bespokelabs.nimble import Nimble

client = Nimble()
response = client.with_raw_response.system_one(
    state="Please refund the duplicate payment.",
    questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
)
print(response.headers.get('X-My-Header'))

result = response.parse()  # get the object that `nimble.system_one()` would have returned
print(result.nouls["refund"].noul)

These methods return an APIResponse object.

The async client returns an AsyncAPIResponse with the same structure, the only difference being awaitable methods for reading the response content.

.with_streaming_response

The above interface eagerly reads the full response body when you make the request, which may not always be what you want.

To stream the response body, use .with_streaming_response instead, which requires a context manager and only reads the response body once you call .read(), .text(), .json(), .iter_bytes(), .iter_text(), .iter_lines() or .parse(). In the async client, these are async methods.

with client.with_streaming_response.system_one(
    state="Please refund the duplicate payment.",
    questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
) as response:
    print(response.headers.get("X-My-Header"))

    for line in response.iter_lines():
        print(line)

The context manager is required so that the response will reliably be closed.

Making custom/undocumented requests

This library is typed for convenient access to the documented API.

If you need to access undocumented endpoints, params, or response properties, the library can still be used.

Undocumented endpoints

To make requests to undocumented endpoints, you can make requests using client.get, client.post, and other http verbs. Options on the client will be respected (such as retries) will be respected when making this request.

import httpx

response = client.post(
    "/foo",
    cast_to=httpx.Response,
    body={"my_param": True},
)

print(response.headers.get("x-foo"))

Undocumented request params

If you want to explicitly send an extra param, you can do so with the extra_query, extra_body, and extra_headers request options.

Undocumented response properties

To access undocumented response properties, you can access the extra fields like response.unknown_prop. You can also get all the extra fields on the Pydantic model as a dict with response.model_extra.

Configuring the HTTP client

You can directly override the httpx client to customize it for your use case, including:

import httpx
from bespokelabs.nimble import Nimble, DefaultHttpxClient

client = Nimble(
    # Or use the `BESPOKE_LABS_BASE_URL` env var
    base_url="http://my.test.server.example.com:8083",
    http_client=DefaultHttpxClient(
        proxy="http://my.test.proxy.example.com",
        transport=httpx.HTTPTransport(local_address="0.0.0.0"),
    ),
)

You can also customize the client on a per-request basis by using with_options():

client.with_options(http_client=DefaultHttpxClient(...))

Managing HTTP resources

By default the library closes underlying HTTP connections whenever the client is garbage collected. You can manually close the client using the .close() method if desired, or with a context manager that closes when exiting.

from bespokelabs.nimble import Nimble

with Nimble() as client:
  # make requests here
  ...

# HTTP client is now closed

Versioning

This package generally follows SemVer conventions, though certain backwards-incompatible changes may be released as minor versions:

  1. Changes that only affect static types, without breaking runtime behavior.
  2. Changes to library internals which are technically public but not intended or documented for external use. (Please open a GitHub issue to let us know if you are relying on such internals.)
  3. Changes that we do not expect to impact the vast majority of users in practice.

We take backwards-compatibility seriously and work hard to ensure you can rely on a smooth upgrade experience.

We are keen for your feedback; please open an issue with questions, bugs, or suggestions.

Determining the installed version

If you've upgraded to the latest version but aren't seeing any new features you were expecting then your python environment is likely still using an older version.

You can determine the version that is being used at runtime with:

import bespokelabs.nimble
print(bespokelabs.nimble.__version__)

Requirements

Python 3.8 or higher.

Contributing

See the contributing documentation.

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

0 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages