Local AI

Research with a local model and a self-hosted search engine

A 27B open-weight model paired with a self-hosted metasearch engine on one workstation, tested for private, local research. What was measured, where it slowed down and what changed.

Pixedi AI Lab
  • 6 min read

SummaryUnder a minute

The short version

A local 27B model paired with a self-hosted metasearch engine runs on one workstation. After the search sources were tuned, results spread across far more domains, and the official model release beat the compressed one on both accuracy and speed, while GPU memory became the bottleneck. The search engines still see the queries and the IP address, but the vendor middleman no longer does.

Key takeaways

  1. After the search sources were changed, every query returned 73-93 hits from 43-63 unique domains in 2.4-10 s.
  2. The official 27B release scored 5/5 on all three passes of a hard task set at about 107 tok/s, while the uncensored quantized version managed 3/5 and 4/5 at a much slower 45 tok/s.
  3. The earlier default model dropped from 46 to 5 tok/s when the image tool held GPU memory, so the assistant now hands memory back and forth on purpose.
  4. A context of 32768 overflowed the working memory and dropped speed to 7 tok/s, so the setting is 16384.
  5. The runtime's built-in web search is off because it routed every query through the vendor's server.
A desktop workstation tower beside a wooden desk in a quiet home office lit by afternoon sun.

The question

Many research questions in technical work are small and specific, like the exact settings a new video model needs in a local image tool. The test here is whether a language model running on a single workstation, connected to a self-hosted search engine, can handle that level of detail well enough to be useful.

The second question is where the queries go. The runtime serving local models has a built-in search that routes everything through the vendor's server, where queries can be seen, tied to an account, counted against a quota, and made dependent on an internet connection and a login. The aim was to see how much of that path can be removed, and to be clear about the parts that can't.

The setup

Everything runs on one workstation with a single consumer GPU that has 24 GB of memory, set up in early September 2026.

The metasearch engine runs in a container on the same machine, bound to the loopback interface only, so nothing else on the network can reach it, and it restarts automatically with the app. It takes a query, sends it to several public engines and merges the results into one list, with no account and no history. Two settings caused trouble. The built-in rate limiter had to be turned off because it kept answering the local assistant with "Too Many Requests" errors, which is safe only because the port is not exposed. The assistant also crashes unless the engine is set to return machine-readable output, so that flag has to stay on.

The second piece is a 27B open-weight model running locally inside a small chat assistant with five tools: search, read the text of a web page, build an image workflow (a saved recipe of steps the image tool follows to make a picture or a clip), start a workflow, and check what the image tool is currently working on. It can't read or write files or run system commands, and that limit is deliberate. It starts an image generation only on an explicit request, and in testing it never started one on its own. A default search pulls 8 results, and a deeper mode downloads the first 4 pages and reads their full text before answering, which is slower but catches details that never appear in a search snippet. Each claim in the answer carries a source number, like [1] or [3], with the URLs listed underneath for checking. Text returned from the web is marked as untrusted, and the assistant is instructed to ignore any instructions inside it, because a web page can contain text written to hijack a model.

The runtime's built-in web search is turned off, because it routed queries through the vendor's server, which is exactly the middleman this setup removes.

Tuning the search sources

The first problem was coverage. The biggest search engine kept filtering out needed results that a Russian index still returned. The metasearch config showed only four active sources, and one of them was broken: it failed to parse the data it received and returned zero results without raising any error. So three engines were doing the work, and since they all drew from the same Western index under the same legal region, they shared the same blind spots.

The broken source was removed and several sources that see the web differently were added, including a Russian index where Western takedown requests don't apply and an independent commercial index built on its own crawler. Chinese and Czech indexes were added too, along with a source that pushes search-optimized pages down and favors small independent sites. That last one needed extra work because its shared public key kept getting rejected with a 429 rate-limit error, so it was switched to its newer interface with a dedicated key.

Measurement

A person at a desk reviewing printed result pages with a pen, next to a coffee cup and a closed laptop.

To check the search tuning, the same queries were run before and after, counting results and the unique domains behind them and timing each search. This measures breadth, so more domains means a wider net, but the count says nothing about whether the best answer lands near the top.

The model was tested on a fixed set of hard tasks where the code it wrote was executed to check the answer, and the full set was repeated three times for each version. Speed is in tokens per second (tok/s), roughly how fast words appear on screen. The limits are plain: the task set is small and everything ran on one machine.

For memory pressure, the GPU versus CPU load split and generation speed were tracked with the image tool holding its memory, then again after that memory was released.

Results

A typical query used to return about 30 results from 20 domains. After the source changes, local search returns 73 to 93 results from 43 to 63 unique domains, in 2.4 to 10 seconds.

On the model side, the first version tried was an uncensored quantized build. Quantized means the weights are compressed so the model takes less memory, and uncensored means a version with the usual refusals removed. On the hard task set it scored 3/5, 4/5 and 4/5 across the three passes, at about 45 tok/s. The official 27B release scored 5/5 in all three passes at about 107 tok/s. It became the default on September 8, 2026, with the uncensored version kept as a one-setting switch for the occasional job that needs it.

Generation speed of the two model versions (current default vs. uncensored quant)tok/s
  • Official 27B release107 tok/s
  • Uncensored quant45 tok/s

Source: Local benchmark on one workstation

The memory and context tests date from September 7, 2026, just before the default changed, so they were run on the older model. The 46 tok/s baseline belongs to that model and shouldn't be compared with the 107 tok/s figure above.

Close-up of a graphics card and memory modules inside an open desktop computer case.

Memory was the biggest surprise, because the image tool and the language model compete for the same 24 GB. With the image tool holding its share, the model's load shifted to 71% CPU and 29% GPU, and generation speed fell from 46 to 5 tok/s. At that pace, waiting for a search answer is impractical.

Earlier default model: speed with and without the image tool holding GPU memorytok/s
  • GPU memory free46 tok/s
  • Image tool holding memory5 tok/s

Source: Local benchmark on one workstation

The assistant now manages memory in both directions. On startup it asks the image tool to release its memory, unless an image is being generated, so work in progress isn't broken. When asked to start an image generation, it unloads itself so the image tool gets the whole GPU, and it loads back in about 10 seconds on the next question. A short chat command frees memory by hand when needed.

Context length causes the same problem in a different place. The model keeps its working memory in the KV cache, and a context setting of 32768 overflowed that cache, which dropped speed from 46 to 7 tok/s. The setting is 16384, where it stays fast.

Earlier default model: speed at two context settingstok/s
  • Context 1638446 tok/s
  • Context 327687 tok/s

Source: Local benchmark on one workstation

Takeaways

One source in the list returned zero results without raising an error, so after any change to the source list, check what each source actually returns.

The 24 GB of GPU memory is the real budget, because the image tool and the model share the same pool, and the model goes from working fine to barely moving the moment the other one takes its share. The CPU and GPU load split is the signal for when that is happening, and a larger context setting looks like an easy improvement while triggering the same kind of slowdown as the memory conflict.

Test candidate models on the actual tasks they will do. That test showed the official release beating the compressed one on both accuracy and speed. A stronger evaluation would also grade the quality of search answers, since counting results says nothing about whether any of them are good.

A desk at dusk with a lamp, notebook, mug and keyboard after a day of work.

Limits in daily use

The setup fits research that needs search and page reading, plus building image workflows on the same machine. It is weaker than a large hosted model at long reasoning and careful writing style, so those tasks are better sent elsewhere. If a request branches too much, it stops after 8 steps and says so. It keeps only the last 24 messages of a conversation and cuts any single tool result at 6000 characters, so in long sessions the older turns fall away, and a fresh conversation works better when the topic changes.

Ask AI about this AI Lab note

Opens the assistant in a new tab with this page as the source.

Keep reading

Want this handled for your business?

Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.