best_engine_ai_helper.score module

score — select the best model for the current hardware.

Given the detected memory pool and the merged catalog, this module picks the highest-scoring model that fits within a safety headroom. The algorithm is intentionally simple: filter, then sort by benchmark score, then take the max.

Memory priority mirrors what Ollama uses at runtime:

unified_gb (Apple Silicon) > vram_gb (discrete GPU) > ram_gb * 0.5 (CPU)

The 0.5 factor for CPU RAM is conservative; a model loader competes with the OS, background daemons, and the inference server itself for RAM.

Author

Warith Harchaoui <warith.harchaoui@deraison.ai>

best_engine_ai_helper.score.effective_budget(hw, headroom=0.85)[source]

Compute the memory budget (GB) a model may occupy at run time.

The budget is the accelerator’s usable memory pool, scaled by an extra headroom margin left for the operating system, your own application, and KV-cache growth as context fills. On Apple Silicon the usable pool is not the whole unified memory: Metal caps GPU allocations at about 66% of the pool at or below 36 GB and about 75% above it, beyond which inference spills to CPU. Compare a catalog entry’s ram_gb (already a peak-inference estimate, weights plus a moderate KV cache) against this budget.

Parameters:
  • hw (dict[str, float | None]) – Output of detect.available_memory(). Expected keys: unified_gb, vram_gb, ram_gb.

  • headroom (float) – Extra safety fraction applied on top of the accelerator cap, reserving room for the OS, the caller’s workload, and KV growth. Default 0.85.

Returns:

Effective memory budget in GB.

Return type:

float

Examples

>>> effective_budget({'unified_gb': 96.0, 'vram_gb': None, 'ram_gb': 96.0})
61.2
>>> effective_budget({'unified_gb': None, 'vram_gb': 24.0, 'ram_gb': 64.0})
18.768
best_engine_ai_helper.score.estimated_tokens_per_second(entry, bandwidth_gbs)[source]

Estimate local decode throughput (tokens/s) for a model on this machine.

Token generation is memory-bandwidth bound: each new token requires reading the model’s active weights from memory once, so the ceiling is bandwidth / model_bytes. The estimate derates that ceiling by _DECODE_EFFICIENCY to reflect KV-cache reads, kernel overhead, and sampling. It describes steady-state generation, not the compute-bound prefill of a long prompt.

Parameters:
  • entry (dict[str, Any]) – Catalog entry; uses ram_gb as the active-model size proxy.

  • bandwidth_gbs (float or None) – Memory bandwidth in GB/s from detect.compute_profile(). When None (unknown hardware) the estimate is not computable and None is returned.

Returns:

Estimated tokens/s, rounded to one decimal, or None when bandwidth or model size is unavailable.

Return type:

float or None

best_engine_ai_helper.score.rank(hw, catalog, kind, headroom=0.85, application=None)[source]

Return all candidates sorted by benchmark score, annotated with fit status.

Each entry in the result gets a _fits key (bool) indicating whether it fits within the effective budget. The top entry is identical to what select() would return.

Parameters:
  • hw (dict[str, float | None]) – Output of detect.available_memory().

  • catalog (list[dict[str, Any]]) – Merged model entries from catalog.load_catalog().

  • kind ({'llm', 'vlm'}) – The inference kind to filter and rank by.

  • headroom (float) – Extra safety on top of the accelerator cap. Default 0.85.

  • application (str or None) – Optional use-case keyword ("code", "math", "ocr", "vision", "chat", "generalist"). Drives which benchmark column is used for ranking.

Returns:

Candidates sorted descending by benchmark score, each with _fits.

Return type:

list[dict[str, Any]]

Examples

>>> catalog = [{'id': 'v', 'kind': 'vlm', 'ram_gb': 9.0,
...             'benchmarks': {'vision': 80}}]
>>> rank({'unified_gb': 96.0, 'vram_gb': None, 'ram_gb': 96.0}, catalog, 'vlm')[0]['_fits']
True
best_engine_ai_helper.score.select(hw, catalog, kind, headroom=0.85, application=None)[source]

Pick the best-scoring model that fits in available memory.

Candidates for a ‘vlm’ selection include both VLMs and LLMs with vision capability (kind == ‘vlm’). Candidates for a ‘llm’ selection include text-only LLMs and VLMs (since a VLM handles text-only prompts equally).

Selection order: 1. Filter: keep entries whose ram_gb fits within the effective budget. 2. Rank: structured-output capability first (a model that cannot honour

Ollama structured output is never chosen over one that can), then benchmark score (application-specific if given, else vision for VLM or general for LLM). This matches rank(), so rank(...)[0] and select(...) agree.

  1. Last resort: if nothing fits, return the smallest model in the catalog rather than raising; the caller decides whether to warn the user.

Parameters:
  • hw (dict[str, float | None]) – Output of detect.available_memory().

  • catalog (list[dict[str, Any]]) – Merged model entries from catalog.load_catalog().

  • kind ({'llm', 'vlm'}) – The type of model to select.

  • headroom (float) – Extra safety on top of the accelerator cap. Default 0.85.

  • application (str or None) – Optional use-case keyword ("code", "math", "ocr", "vision", "chat", "generalist"). Selects the benchmark axis used for scoring. None uses the default kind-based rule.

Returns:

The selected catalog entry.

Return type:

dict[str, Any]

Raises:

ValueError – If the catalog is empty.

Examples

>>> hw = {'unified_gb': 96.0, 'vram_gb': None, 'ram_gb': 96.0}
>>> catalog = [{'id': 'v', 'kind': 'vlm', 'ram_gb': 9.0,
...             'benchmarks': {'vision': 80}}]
>>> select(hw, catalog, kind='vlm')['id']
'v'