The Bespoke Nimble Python library provides convenient access to the Bespoke Labs REST API from any Python 3.8+ application. The library includes type definitions for all request params and response fields, and offers both synchronous and asynchronous clients powered by httpx.
It is generated with Stainless.
The REST API documentation can be found on docs.bespokelabs.ai. The full API of this library can be found in api.md.
# install or upgrade from PyPI
pip install --upgrade bespokelabs-nimbleThis checkout prepares the next bespokelabs-nimble release. The commands below
apply once that release is published; PyPI's existing 0.1.0 is the older
standalone SDK. To try this change before release, install this checkout with
pip install . in a fresh virtual environment.
Replace the dependency bespokelabs with bespokelabs-nimble and update imports:
# Before
from bespokelabs import BespokeLabs, AsyncBespokeLabs
# After
from bespokelabs.nimble import Nimble, AsyncNimbleTypes and exceptions also move under bespokelabs.nimble, for example
from bespokelabs.nimble.types.nimble import Question and
from bespokelabs.nimble import APIError.
Use client.system_one(...) for Nimble and client.factcheck(...) for batched fact
checking. BespokeLabs / AsyncBespokeLabs remain aliases of Nimble / AsyncNimble.
The previous client.nimble.system_one(...) and
client.minicheck.factcheck.create(...) calls remain available.
Use a fresh virtual environment when migrating an environment that also has
Curator or Sandbox installed. The old bespokelabs distribution shares its root
bespokelabs/__init__.py with those packages; uninstalling it can remove that
shared file. This SDK now owns only bespokelabs/nimble/ and does not export
clients from the shared root.
Import Nimble / AsyncNimble from bespokelabs.nimble.
The direct client.system_one(...) method remains available, but authentication
and question construction differ from the standalone SDK.
Use question dictionaries as in the examples below instead of the standalone
Noul, Choice, and Score constructors.
The client requires BESPOKE_API_KEY (or api_key=) and defaults to the hosted
Bespoke gateway. Custom gateways use BESPOKE_LABS_BASE_URL (or base_url=)
and must expose /v1/nimble/systemone; a standalone model server exposing
/v1/systemone is not interchangeable. health(), models(), and limits()
from the standalone SDK are not provided by this client. Applications that
still need that interface should keep bespokelabs-nimble==0.1.0 pinned until
they migrate.
Set BESPOKE_API_KEY in your environment:
export BESPOKE_API_KEY="your-api-key"The SDK reads this variable automatically when you create a client:
from bespokelabs.nimble import Nimble
client = Nimble()You can also pass a key explicitly with api_key:
from bespokelabs.nimble import Nimble
client = Nimble(api_key="your-api-key")The same options work with AsyncNimble. An explicit api_key takes
precedence over BESPOKE_API_KEY. Keep real keys out of source control.
If you keep credentials in a .env file, install
python-dotenv and load the file before
creating the client; the SDK does not load .env files itself:
from dotenv import load_dotenv
from bespokelabs.nimble import Nimble
load_dotenv() # Loads BESPOKE_API_KEY from .env into the environment.
client = Nimble()Replace auth_token= with api_key= when constructing a client or calling
copy() or with_options(). Replace client.auth_token with client.api_key.
The BESPOKE_API_KEY environment variable and HTTP authentication header are
unchanged, so clients configured solely through the environment need no changes.
The full API of this library can be found in api.md.
from bespokelabs import nimble
client = nimble.Nimble() # Reads BESPOKE_API_KEY from the environment.
result = client.system_one(
state="Please refund the duplicate payment.",
questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
)
print(result.nouls["refund"].noul)Use Nimble to answer structured questions about text with Noul, Choice, and Score outputs:
from bespokelabs.nimble import Nimble
with Nimble() as client:
result = client.system_one(
state="Please refund the duplicate payment.",
questions={
"refund": {"type": "noul", "instructions": "Does the customer request a refund?"},
"department": {
"type": "choice",
"instructions": "Which department should handle this request?",
"criteria": {"billing": "Payments and refunds", "technical": "Software bugs"},
},
"urgency": {
"type": "score",
"instructions": "Assess operational urgency.",
"criteria": ["Service works normally", "Partial outage", "Complete outage"],
},
},
)
print(result.nouls["refund"].noul)
print(result.choices["department"].choice)
print(result.scores["urgency"].score)The default client sends POST /v1/nimble/systemone to https://api.bespokelabs.ai.
For a custom deployment, set base_url or BESPOKE_LABS_BASE_URL to a gateway exposing
the same route. This route differs from a standalone Nimble model server's /v1/systemone.
A custom gateway without the route returns 404; one without a configured model server returns 503.
Noul returns the probability of true. Choice returns a selected option and the probability
of each option. Score returns the expected zero-based rubric index: for three levels, its
range is 0–2. Confidence measures distribution concentration, not calibrated correctness.
There are at most 64 questions per request and 2–26 candidates per Choice or Score.
The default model is nimble-latest; pass model= to select another server-supported model.
Answers are typed and available through result.answers, or through result.nouls,
result.choices, and result.scores. Question types are exported from
bespokelabs.nimble.types.nimble.
The same resource is available on AsyncNimble:
from bespokelabs.nimble import AsyncNimble
async def check_refund():
async with AsyncNimble() as client:
result = await client.system_one(
state="Please refund the duplicate payment.",
questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
)
return result.nouls["refund"].noulThe standard SDK options work for Nimble, including with_options, retries, per-call timeouts,
client.with_raw_response.system_one(...), and
client.with_streaming_response.system_one(...). Streaming here controls HTTP body
reading; it does not produce incremental model answers. Increase the request timeout for
slow responses (for example, timeout=180.0). A longer timeout does not make a server wait
when it immediately returns an overload or startup error such as 503 or 529; those responses
use the SDK's configured retry policy.
from bespokelabs import nimble
with nimble.Nimble() as client:
response = client.factcheck(
context="Paris is the capital of France.",
claims=["France's capital is Paris.", "France's capital is London."],
effort="medium",
)
for result in response.results:
print(result.claim, result.support_prob)factcheck accepts 1–64 non-empty claims, each up to 4,000 characters, against
up to 400,000 characters of context. It sends one request to
/v1/nimble/factcheck. Results preserve input order and duplicate claims.
support_prob is the probability that the context supports the whole claim;
it is not a calibrated guarantee of truth. supported means the score is above 0.5.
All three effort values are supported:
"low": score with the small fact-checking model."medium"(default): recheck uncertain claims with a larger model."high": recheck uncertain claims with additional judging and reasoning models.
split_claims=True is the default. Each sentence is checked separately and the
lowest support score is returned for the original claim. Set it to False to
check each claim whole. Medium and high may return escalated and per-model
scores for each claim. If the requested tier cannot finish because its additional
models are starting, the API returns a retryable 503 and does not charge the request.
The response includes model, effort, usage, and request_id. Billable input
counts the original context once and each original claim once, using the service's
shared tokenizer. Internal splitting, chunking and reasoning do not multiply that
count; output_tokens is zero because the response contains scores, not generated
text. The selected effort determines the input-token rate.
Use await client.factcheck(...) with AsyncNimble. Both clients accept
extra_headers, extra_query, and timeout. The effort selects the model;
model is not an argument to this method. The legacy
client.minicheck.factcheck.create(claim=..., context=...) is unchanged.
Simply import AsyncNimble instead of Nimble and use await with each API call:
import asyncio
from bespokelabs.nimble import AsyncNimble
client = AsyncNimble() # Reads BESPOKE_API_KEY from the environment.
async def main() -> None:
result = await client.system_one(
state="Please refund the duplicate payment.",
questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
)
print(result.nouls["refund"].noul)
asyncio.run(main())Functionality between the synchronous and asynchronous clients is otherwise identical.
Nested request parameters are TypedDicts. Responses are Pydantic models which also provide helper methods for things like:
- Serializing back into JSON,
model.to_json() - Converting to a dictionary,
model.to_dict()
Typed requests and responses provide autocomplete and documentation within your editor. If you would like to see type errors in VS Code to help catch bugs earlier, set python.analysis.typeCheckingMode to basic.
When the library is unable to connect to the API (for example, due to network connection problems or a timeout), a subclass of bespokelabs.nimble.APIConnectionError is raised.
When the API returns a non-success status code (that is, 4xx or 5xx
response), a subclass of bespokelabs.nimble.APIStatusError is raised, containing status_code and response properties.
All errors inherit from bespokelabs.nimble.APIError.
import bespokelabs.nimble
from bespokelabs.nimble import Nimble
client = Nimble()
try:
client.system_one(
state="Please refund the duplicate payment.",
questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
)
except bespokelabs.nimble.APIConnectionError as e:
print("The server could not be reached")
print(e.__cause__) # an underlying Exception, likely raised within httpx.
except bespokelabs.nimble.RateLimitError as e:
print("A 429 status code was received; we should back off a bit.")
except bespokelabs.nimble.APIStatusError as e:
print("Another non-200-range status code was received")
print(e.status_code)
print(e.response)Error codes are as followed:
| Status Code | Error Type |
|---|---|
| 400 | BadRequestError |
| 401 | AuthenticationError |
| 403 | PermissionDeniedError |
| 404 | NotFoundError |
| 422 | UnprocessableEntityError |
| 429 | RateLimitError |
| >=500 | InternalServerError |
| N/A | APIConnectionError |
Certain errors are automatically retried 2 times by default, with a short exponential backoff. Connection errors (for example, due to a network connectivity problem), 408 Request Timeout, 409 Conflict, 429 Rate Limit, and >=500 Internal errors are all retried by default.
You can use the max_retries option to configure or disable retry settings:
from bespokelabs.nimble import Nimble
# Configure the default for all requests:
client = Nimble(
# default is 2
max_retries=0,
)
# Or, configure per-request:
client.with_options(max_retries=5).nimble.system_one(
state="Please refund the duplicate payment.",
questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
)By default requests time out after 1 minute. You can configure this with a timeout option,
which accepts a float or an httpx.Timeout object:
import httpx
from bespokelabs.nimble import Nimble
# Configure the default for all requests:
client = Nimble(
# 20 seconds (default is 1 minute)
timeout=20.0,
)
# More granular control:
client = Nimble(
timeout=httpx.Timeout(60.0, read=5.0, write=10.0, connect=2.0),
)
# Override per-request:
client.with_options(timeout=5.0).nimble.system_one(
state="Please refund the duplicate payment.",
questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
)On timeout, an APITimeoutError is thrown.
Note that requests that time out are retried twice by default.
We use the standard library logging module.
You can enable logging by setting the environment variable BESPOKE_LABS_LOG to info.
$ export BESPOKE_LABS_LOG=infoOr to debug for more verbose logging.
In an API response, a field may be explicitly null, or missing entirely; in either case, its value is None in this library. You can differentiate the two cases with .model_fields_set:
if response.my_field is None:
if 'my_field' not in response.model_fields_set:
print('Got json like {}, without a "my_field" key present at all.')
else:
print('Got json like {"my_field": null}.')The "raw" Response object can be accessed by prefixing .with_raw_response. to any HTTP method call, e.g.,
from bespokelabs.nimble import Nimble
client = Nimble()
response = client.with_raw_response.system_one(
state="Please refund the duplicate payment.",
questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
)
print(response.headers.get('X-My-Header'))
result = response.parse() # get the object that `nimble.system_one()` would have returned
print(result.nouls["refund"].noul)These methods return an APIResponse object.
The async client returns an AsyncAPIResponse with the same structure, the only difference being awaitable methods for reading the response content.
The above interface eagerly reads the full response body when you make the request, which may not always be what you want.
To stream the response body, use .with_streaming_response instead, which requires a context manager and only reads the response body once you call .read(), .text(), .json(), .iter_bytes(), .iter_text(), .iter_lines() or .parse(). In the async client, these are async methods.
with client.with_streaming_response.system_one(
state="Please refund the duplicate payment.",
questions={"refund": {"type": "noul", "instructions": "Refund requested?"}},
) as response:
print(response.headers.get("X-My-Header"))
for line in response.iter_lines():
print(line)The context manager is required so that the response will reliably be closed.
This library is typed for convenient access to the documented API.
If you need to access undocumented endpoints, params, or response properties, the library can still be used.
To make requests to undocumented endpoints, you can make requests using client.get, client.post, and other
http verbs. Options on the client will be respected (such as retries) will be respected when making this
request.
import httpx
response = client.post(
"/foo",
cast_to=httpx.Response,
body={"my_param": True},
)
print(response.headers.get("x-foo"))If you want to explicitly send an extra param, you can do so with the extra_query, extra_body, and extra_headers request
options.
To access undocumented response properties, you can access the extra fields like response.unknown_prop. You
can also get all the extra fields on the Pydantic model as a dict with
response.model_extra.
You can directly override the httpx client to customize it for your use case, including:
- Support for proxies
- Custom transports
- Additional advanced functionality
import httpx
from bespokelabs.nimble import Nimble, DefaultHttpxClient
client = Nimble(
# Or use the `BESPOKE_LABS_BASE_URL` env var
base_url="http://my.test.server.example.com:8083",
http_client=DefaultHttpxClient(
proxy="http://my.test.proxy.example.com",
transport=httpx.HTTPTransport(local_address="0.0.0.0"),
),
)You can also customize the client on a per-request basis by using with_options():
client.with_options(http_client=DefaultHttpxClient(...))By default the library closes underlying HTTP connections whenever the client is garbage collected. You can manually close the client using the .close() method if desired, or with a context manager that closes when exiting.
from bespokelabs.nimble import Nimble
with Nimble() as client:
# make requests here
...
# HTTP client is now closedThis package generally follows SemVer conventions, though certain backwards-incompatible changes may be released as minor versions:
- Changes that only affect static types, without breaking runtime behavior.
- Changes to library internals which are technically public but not intended or documented for external use. (Please open a GitHub issue to let us know if you are relying on such internals.)
- Changes that we do not expect to impact the vast majority of users in practice.
We take backwards-compatibility seriously and work hard to ensure you can rely on a smooth upgrade experience.
We are keen for your feedback; please open an issue with questions, bugs, or suggestions.
If you've upgraded to the latest version but aren't seeing any new features you were expecting then your python environment is likely still using an older version.
You can determine the version that is being used at runtime with:
import bespokelabs.nimble
print(bespokelabs.nimble.__version__)Python 3.8 or higher.