<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Ia on lo0 — Tech Blog</title><link>https://blog.lo0.es/en/categories/ia/</link><description>Recent content in Ia on lo0 — Tech Blog</description><generator>Hugo -- gohugo.io</generator><language>en</language><lastBuildDate>Tue, 15 Sep 2026 09:30:00 +0200</lastBuildDate><atom:link href="https://blog.lo0.es/en/categories/ia/index.xml" rel="self" type="application/rss+xml"/><item><title>Real ClickHouse capacity in Langfuse v4: 2,782 bytes per observation, and why almost half of that is not data</title><link>https://blog.lo0.es/en/posts/clickhouse-capacity-langfuse-v4-bytes-per-observation/</link><pubDate>Tue, 15 Sep 2026 09:30:00 +0200</pubDate><guid>https://blog.lo0.es/en/posts/clickhouse-capacity-langfuse-v4-bytes-per-observation/</guid><description>&lt;blockquote>
&lt;p>Sixth article in the series on operating self-hosted Langfuse v4. The previous five cover &lt;a href="https://blog.lo0.es/en/posts/langfuse-v4-what-goes-into-a-trace/">what goes into a trace&lt;/a>, &lt;a href="https://blog.lo0.es/en/posts/langfuse-v4-instrumenting-langgraph/">putting LangGraph in front&lt;/a>, &lt;a href="https://blog.lo0.es/en/posts/langfuse-v4-real-agents-odr-gptr/">two well-known agents instrumented&lt;/a>, &lt;a href="https://blog.lo0.es/en/posts/langfuse-v4-migrating-with-no-window/">migrating with no window&lt;/a> and &lt;a href="https://blog.lo0.es/en/posts/langfuse-v4-worker-queues/">the worker queues&lt;/a>. This one measures what all of that takes up. Benchmark run on 15 September 2026 against the event schema of version 4.36.0, identical to that of 4.35.0.&lt;/p>
&lt;/blockquote>
&lt;h2 id="tldr">TL;DR&lt;/h2>
&lt;p>&lt;strong>Langfuse publishes no capacity figure at all.&lt;/strong> Not bytes per row, not events per second, not expected size. I searched their code with every reasonable combination and there is nothing. The only quantified thing is the commercial quotas of their managed service, which say nothing about disk.&lt;/p>
&lt;p>&lt;strong>One observation takes 2,782 bytes of disk.&lt;/strong> Split across the two event tables: 2,347 in the full table and 435 in the listings one. Measured over two hundred thousand observations shaped the way a real agent produces them.&lt;/p>
&lt;p>&lt;strong>41.8 % of the full table&amp;rsquo;s disk is not data, it is indexes.&lt;/strong> Four full-text indexes add up to 187 MiB against 260 MiB of data. The one covering the input costs 523 bytes per observation, almost three quarters of what the column it indexes costs.&lt;/p>
&lt;p>&lt;strong>The materialized view trims the text eight times over and saves eighteen per cent of disk.&lt;/strong> Truncating input and output to two hundred characters leaves the text at 64 MiB of the original 540, and even so the second table adds 18.5 % on top of the first. What survives the trimming is the seventy-eight-column skeleton and a metadata field that is truncated element by element, not as a whole.&lt;/p>
&lt;p>&lt;strong>In the listings table, metadata costs more than input and output together.&lt;/strong> 132.6 bytes per observation against 101.8. And its full-text index adds another 111.9, almost as much as the column.&lt;/p>
&lt;p>&lt;strong>A compression ratio quoted without saying what was compressed is worth nothing.&lt;/strong> My first pass, with repeated text, gave 38 times. With text of an entropy close to that of natural language it gives 3.68. Same schema, same indexes, same version.&lt;/p>
&lt;p>&lt;strong>Generations are 11 % of the observations and 65 % of the bytes.&lt;/strong> Their mean input is 15,655 bytes and the 95th percentile is 55,205, because each one reserialises the full history. It is the quadratic law from the second article in the series, landed on disk.&lt;/p>
&lt;p>&lt;strong>On deletion, the disk does not go down.&lt;/strong> A delete stops showing the rows and leaves the space occupied. A second step is needed, applying the deleted mask, and the cron that does it &lt;strong>ships disabled&lt;/strong>, just like the retention one. On top of that the cleaner works table by table, so the second table keeps the rows the first one no longer shows.&lt;/p>
&lt;p>&lt;strong>And automatic retention is a paid feature when self-hosted.&lt;/strong> The &lt;code>data-retention&lt;/code> entitlement is not in the open edition. Without it, the event tables grow without limit: they carry no expiry clause.&lt;/p>
&lt;h2 id="you-are-here-the-sizing-you-cannot-do-by-reading">You are here: the sizing you cannot do by reading&lt;/h2>
&lt;p>The previous five articles describe how version 4 works and how to put it to work. The moment arrives to ask for disk, and there the series runs out of sources.&lt;/p>
&lt;p>The official documentation gives a container sizing: two CPUs and eight gibibytes for ClickHouse, in Kubernetes eight of request and sixteen of limit, a hundred gibibytes of disk and three minimum replicas. It is a reasonable starting point and it does not answer the only question that matters, which is how long that disk lasts.&lt;/p>
&lt;p>I went looking for it in the code, which is what this series does. I searched for events per second, rows per observation, bytes per row, capacity, benchmark, sizing, across the whole repository and its documentation. &lt;strong>There is no figure at all.&lt;/strong> The closest thing is the ingestion quotas of their managed service, which are commercial API limits, and some working notes on a synthetic load generator that refer to the tables of the previous version and bring no results, only parameters.&lt;/p>
&lt;p>So the only honest way out was to measure it. This article is that measurement, with the whole method laid out so that anyone can repeat it and contrast it with their own load.&lt;/p>
&lt;h2 id="the-analogy-the-archive-and-the-index-card">The analogy: the archive and the index card&lt;/h2>
&lt;p>A municipal archive keeps case files. Each file has its folder with everything inside, and on top of that the archivist fills in an index card with the header data and the first lines of the subject, so that the file can be searched without pulling out the folder.&lt;/p>
&lt;p>Intuition says the cards take up nothing compared with the folders. And that is true if you look only at the paper. It stops being true when the archivist, in order to be able to search by any word, also builds a term file: a box with one card per distinct word appearing in any case file, pointing to where it occurs. That box grows with the vocabulary, not with the number of case files, and ends up taking almost as much room as the case files themselves.&lt;/p>
&lt;p>In this setup the folders are the full table, the cards are the listings table, and the term file is the full-text indexes. The article measures the three things separately, because whoever sizes by looking only at the folders falls short by almost half.&lt;/p>
&lt;h2 id="the-benchmark">The benchmark&lt;/h2>
&lt;p>I want this to be reproducible, so the method goes in whole.&lt;/p>
&lt;p>&lt;strong>The schema is the real one.&lt;/strong> I extracted the ClickHouse migration DDL from the repository, in its single-instance variant. There are not three migrations but six: the three that create the two tables and the materialized view, plus one that adds three ingestion attribution columns, another that adds an n-gram index over the metadata, and another that adds three evaluation columns and two indexes. The result is &lt;strong>78 columns per table, eleven skip indexes in the full table and fourteen in the listings one&lt;/strong>.&lt;/p>
&lt;p>I verified that this DDL is identical between the version I have pinned across the series and the current head of the repository: the diff of the migrations directory is empty.&lt;/p>
&lt;p>&lt;strong>The engine is ClickHouse 26.9.1.&lt;/strong> The minimum version 4 demands is 25.12, because of the text indexes the schema needs. I am using a newer version, and that is a difference from an installation that sticks to the minimum.&lt;/p>
&lt;p>&lt;strong>The data is synthetic and correctly shaped.&lt;/strong> The benchmark does not invent the rows: the columns and their typical values come from the worker function that builds the record. I write only into the full table, which is what the worker does, and let the materialized view populate the second one. That reproduces the real write amplification.&lt;/p>
&lt;p>The topology of each trace reproduces the law we measured in the second article of the series: four fixed observations plus five for each tool call. The number of calls per turn comes from a distribution with a low median and a long tail, capped at fifteen. Generations carry the full message history, which grows with the turn, because that is exactly what the instrumentation does and it is the origin of the quadratic growth in bytes.&lt;/p>
&lt;p>Result: 200,012 observations across 22,563 traces, that is &lt;strong>8.86 observations per trace&lt;/strong>.&lt;/p>
&lt;p>&lt;strong>And the entropy of the text is calibrated, which is the part that almost ruined the benchmark.&lt;/strong> My first pass used a repeated paragraph as filler. With that text, the full table compressed 38.1 times and the figures came out beautiful and false. A repeated lorem ipsum is the best possible input for a compressor. I redid the generator with a wide vocabulary sampled from a Zipf distribution and fifteen per cent high-entropy identifiers, calibrated to give the ratio that natural language mixed with identifiers gives.&lt;/p>
&lt;p>With that text, the same table compresses 3.68 times. &lt;strong>Ten times less.&lt;/strong> Any ClickHouse capacity figure that does not say what was compressed is noise, and mine would be too without this paragraph.&lt;/p>
&lt;h2 id="what-one-observation-takes-up">What one observation takes up&lt;/h2>
&lt;p>The two tables, after forcing the final merge:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>&lt;/th>
&lt;th>events_full&lt;/th>
&lt;th>events_core&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Rows&lt;/td>
&lt;td>200,012&lt;/td>
&lt;td>200,012&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Uncompressed&lt;/td>
&lt;td>957.87 MiB&lt;/td>
&lt;td>328.73 MiB&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Compressed&lt;/td>
&lt;td>260.29 MiB&lt;/td>
&lt;td>60.34 MiB&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Ratio&lt;/td>
&lt;td>3.68&lt;/td>
&lt;td>5.45&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>On disk&lt;/td>
&lt;td>447.68 MiB&lt;/td>
&lt;td>82.92 MiB&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Raw per observation&lt;/td>
&lt;td>5,022 B&lt;/td>
&lt;td>1,723 B&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Compressed per observation&lt;/td>
&lt;td>1,364.6 B&lt;/td>
&lt;td>316.3 B&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;strong>On disk per observation&lt;/strong>&lt;/td>
&lt;td>&lt;strong>2,347 B&lt;/strong>&lt;/td>
&lt;td>&lt;strong>435 B&lt;/strong>&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>The figure to take away is the last row added up: &lt;strong>2,782 bytes of disk per observation&lt;/strong>, counting both tables, with one replica and no retention.&lt;/p>
&lt;p>And it is worth looking at the distance between the last two rows. The compressed data of the full table is 1,364.6 bytes per observation, but on disk it takes 2,347. The difference, 982 bytes per observation, is the next section.&lt;/p>
&lt;h2 id="where-the-bytes-go">Where the bytes go&lt;/h2>
&lt;p>Broken down by column, with what each one costs per observation:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Column&lt;/th>
&lt;th>events_full&lt;/th>
&lt;th>events_core&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>input&lt;/code>&lt;/td>
&lt;td>717.5 B&lt;/td>
&lt;td>61.5 B&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>metadata_values&lt;/code>&lt;/td>
&lt;td>389.0 B&lt;/td>
&lt;td>132.6 B&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>output&lt;/code>&lt;/td>
&lt;td>174.7 B&lt;/td>
&lt;td>40.3 B&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>tool_definitions&lt;/code>&lt;/td>
&lt;td>28.0 B&lt;/td>
&lt;td>28.1 B&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>blob_storage_file_path&lt;/code>&lt;/td>
&lt;td>12.6 B&lt;/td>
&lt;td>12.5 B&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>span_id&lt;/code>&lt;/td>
&lt;td>8.1 B&lt;/td>
&lt;td>8.0 B&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>The other 67 columns together&lt;/td>
&lt;td>17.3 B&lt;/td>
&lt;td>13.3 B&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>Three things jump out.&lt;/p>
&lt;p>The first is that &lt;strong>the skeleton is cheap&lt;/strong>. Sixty-seven columns, most of them empty in a normal observation, cost seventeen bytes between them all. ClickHouse compresses a column of empty strings down to almost nothing. Anyone who feared that the thirteen experiment columns or the telemetry ones would weigh can stop fearing it.&lt;/p>
&lt;p>The second is that in the full table &lt;strong>the input is half of everything&lt;/strong>. 717.5 of 1,364.6 compressed bytes. It is the direct consequence of each generation reserialising the history.&lt;/p>
&lt;p>The third is the one I was not expecting. In the listings table, &lt;strong>metadata costs more than input and output together&lt;/strong>: 132.6 against 101.8. The reason lies in how the materialized view truncates. It trims input and output to two hundred characters each, but it trims the metadata &lt;strong>element by element&lt;/strong>: an array with forty values can survive with eight thousand characters. The table that exists in order to be light fills up with the one thing its truncation does not control.&lt;/p>
&lt;h3 id="the-split-by-observation-type">The split by observation type&lt;/h3>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Type&lt;/th>
&lt;th>Observations&lt;/th>
&lt;th>Mean input&lt;/th>
&lt;th>Input p95&lt;/th>
&lt;th>Mean output&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>GENERATION&lt;/td>
&lt;td>21,952&lt;/td>
&lt;td>15,655 B&lt;/td>
&lt;td>55,205 B&lt;/td>
&lt;td>1,077 B&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>AGENT&lt;/td>
&lt;td>22,563&lt;/td>
&lt;td>2,585 B&lt;/td>
&lt;td>6,013 B&lt;/td>
&lt;td>1,300 B&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>TOOL&lt;/td>
&lt;td>21,952&lt;/td>
&lt;td>319 B&lt;/td>
&lt;td>319 B&lt;/td>
&lt;td>2,008 B&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>CHAIN&lt;/td>
&lt;td>67,689&lt;/td>
&lt;td>413 B&lt;/td>
&lt;td>413 B&lt;/td>
&lt;td>20 B&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>SPAN&lt;/td>
&lt;td>65,856&lt;/td>
&lt;td>272 B&lt;/td>
&lt;td>272 B&lt;/td>
&lt;td>206 B&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>Generations are 11 % of the observations and contribute 65 % of the input and output bytes. The 95th percentile of their input is 55 kilobytes, in a benchmark where no conversation goes beyond fifteen tool calls.&lt;/p>
&lt;p>This has a sizing consequence worth saying out loud: &lt;strong>the size of the database is not determined by the number of requests, it is determined by the depth of the turns&lt;/strong>. Doubling the users doubles the disk. Doubling the tool calls per turn multiplies it by rather more than two.&lt;/p>
&lt;h2 id="the-indexes-which-are-almost-half">The indexes, which are almost half&lt;/h2>
&lt;p>Here is the finding that changes a sizing exercise the most.&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>&lt;/th>
&lt;th>Data&lt;/th>
&lt;th>Indexes&lt;/th>
&lt;th>% of disk in indexes&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>events_full&lt;/td>
&lt;td>260.29 MiB&lt;/td>
&lt;td>187.31 MiB&lt;/td>
&lt;td>&lt;strong>41.8 %&lt;/strong>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>events_core&lt;/td>
&lt;td>60.34 MiB&lt;/td>
&lt;td>22.49 MiB&lt;/td>
&lt;td>&lt;strong>27.1 %&lt;/strong>&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>The marks and the primary key are irrelevant: thirty-two kilobytes and three hundred bytes respectively, with the granule size of 64 mebibytes the full table has declared.&lt;/p>
&lt;p>Broken down, and ordered by what they cost:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Index&lt;/th>
&lt;th>Table&lt;/th>
&lt;th>Bytes per observation&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>idx_fts_input_low&lt;/code>&lt;/td>
&lt;td>events_full&lt;/td>
&lt;td>523.1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>idx_fts_metadata_values&lt;/code>&lt;/td>
&lt;td>events_full&lt;/td>
&lt;td>319.4&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>idx_fts_output_low&lt;/code>&lt;/td>
&lt;td>events_full&lt;/td>
&lt;td>133.5&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>idx_fts_metadata_values&lt;/code>&lt;/td>
&lt;td>events_core&lt;/td>
&lt;td>111.9&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>idx_fts_metadata_names&lt;/code>&lt;/td>
&lt;td>both&lt;/td>
&lt;td>4.3&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>idx_span_id&lt;/code>&lt;/td>
&lt;td>both&lt;/td>
&lt;td>1.2&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>The remaining Bloom filters&lt;/td>
&lt;td>both&lt;/td>
&lt;td>0.1 each&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>The full-text index over the input costs 523 bytes per observation when the column it indexes costs 717. That is, &lt;strong>indexing the input costs almost three quarters of what storing it costs&lt;/strong>. And the metadata one in the listings table costs 111.9 when its column costs 132.6: eighty-five per cent.&lt;/p>
&lt;p>The Bloom filters, on the other hand, are free as far as sizing goes. The four identifier ones and those for model name, experiment and evaluator add up to less than two bytes per observation between them all.&lt;/p>
&lt;p>This is not a criticism of the design. Those indexes are what make text search in the interface respond, and without them the experience would be another thing. It is sizing information that is published nowhere: &lt;strong>when asking for disk for ClickHouse you have to multiply what the columns say by 1.7.&lt;/strong>&lt;/p>
&lt;h2 id="the-materialized-view-as-an-amplifier">The materialized view as an amplifier&lt;/h2>
&lt;p>The second table exists so that listings do not have to touch the large fields. It truncates input and output to two hundred characters and does without the text indexes over them.&lt;/p>
&lt;p>The trimming works: the benchmark&amp;rsquo;s 540.55 MiB of input and output come down to 64.68 MiB, &lt;strong>eight and a half times less&lt;/strong>.&lt;/p>
&lt;p>And even so the second table adds 18.5 % of disk on top of the first. For the reasons already seen: the skeleton survives, the metadata truncated element by element survives, and its own text index over that metadata costs another 111.9 bytes.&lt;/p>
&lt;p>It is worth keeping in mind because it is counter-intuitive. You would think a table storing two hundred characters of text is an appendix. It is four hundred and thirty-five bytes per observation, almost a fifth of the total, and it is paid always, because the materialized view cannot be switched off without leaving the interface with no listings.&lt;/p>
&lt;h2 id="what-happens-on-deletion">What happens on deletion&lt;/h2>
&lt;p>The retention cron runs a lightweight delete. I reproduced it: I deleted the observations older than a date, which removed 67,395 rows of the 200,012.&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Moment&lt;/th>
&lt;th>Visible rows&lt;/th>
&lt;th>Rows in the parts&lt;/th>
&lt;th>Disk&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Before&lt;/td>
&lt;td>200,012&lt;/td>
&lt;td>200,012&lt;/td>
&lt;td>447.68 MiB&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>After the delete&lt;/td>
&lt;td>132,617&lt;/td>
&lt;td>200,012&lt;/td>
&lt;td>447.68 MiB&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>After applying the mask&lt;/td>
&lt;td>132,617&lt;/td>
&lt;td>132,617&lt;/td>
&lt;td>299.11 MiB&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>The delete frees &lt;strong>not one byte&lt;/strong>. The rows stop being visible and the space stays occupied. The disk only goes down at the second step, applying the deleted mask, which frees 33.2 %.&lt;/p>
&lt;p>And that second step is run by a different cron that &lt;strong>ships disabled&lt;/strong>, just like the retention one. With the default values, whoever switches retention on in the interface will see the traces disappear from the view and will never see the disk go down, until an ordinary merge drags in the affected parts, which can take however long it takes.&lt;/p>
&lt;p>There is a third detail. The cleaner works &lt;strong>table by table&lt;/strong>, with an independent deleter for each one. In my test I deleted in the full table and the listings one kept its 200,012 rows, because the materialized view does not propagate deletes. In a real installation both deleters run, but they are two separate operations that can diverge if one fails.&lt;/p>
&lt;p>On top of this you have to add what the fourth article already covered: &lt;strong>neither of the two tables carries an expiry clause&lt;/strong>. There is no automatic expiry in the engine. The whole lifecycle depends on those crons. And per-project retention requires an enterprise edition entitlement, so in the open edition the only retention possible is the one you write yourself.&lt;/p>
&lt;h2 id="projection">Projection&lt;/h2>
&lt;p>With the measured 2,782 bytes per observation, one replica and no retention:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Observations per day&lt;/th>
&lt;th>Agent turns per day&lt;/th>
&lt;th>Per day&lt;/th>
&lt;th>Per month&lt;/th>
&lt;th>Per year&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>100,000&lt;/td>
&lt;td>~11,300&lt;/td>
&lt;td>0.26 GiB&lt;/td>
&lt;td>7.8 GiB&lt;/td>
&lt;td>94.6 GiB&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>1,000,000&lt;/td>
&lt;td>~112,800&lt;/td>
&lt;td>2.59 GiB&lt;/td>
&lt;td>77.7 GiB&lt;/td>
&lt;td>945.6 GiB&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>10,000,000&lt;/td>
&lt;td>~1,128,000&lt;/td>
&lt;td>25.91 GiB&lt;/td>
&lt;td>777.2 GiB&lt;/td>
&lt;td>9,456 GiB&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>The conversion to turns uses the benchmark&amp;rsquo;s 8.86 observations per trace, which depend on the depth of the agents. With deeper agents, the same turn figure produces more observations.&lt;/p>
&lt;p>And there is still a multiplication left. Langfuse recommends &lt;strong>three minimum replicas&lt;/strong> of ClickHouse in Kubernetes. With one million observations a day and three replicas, a year is &lt;strong>2.8 tebibytes of raw disk&lt;/strong>. The official sizing of a hundred gibibytes per replica covers about thirty-eight days at that rate.&lt;/p>
&lt;p>That is the number that is published nowhere and that has to go into the spreadsheet before committing to a service.&lt;/p>
&lt;h2 id="what-this-benchmark-does-not-measure">What this benchmark does not measure&lt;/h2>
&lt;p>For the sake of honesty, and because anyone repeating this is going to get different numbers:&lt;/p>
&lt;p>The data is synthetic. The shape of the row and the topology of the traces are taken from the code and from the measurements in the previous articles, but the text is generated by a sampler, not by a model. I have calibrated its entropy so that compression resembles that of natural language, and that calibration is the largest source of error in the whole article. A load with a lot of structured JSON will compress better; one with a lot of code or a lot of base64, worse.&lt;/p>
&lt;p>It is a single instance, with no replication. In a replicated installation there is a different engine and there are copies.&lt;/p>
&lt;p>It is ClickHouse 26.9.1, above the 25.12 minimum. The text index implementations have changed between versions and I have not compared.&lt;/p>
&lt;p>It does not measure performance: not ingestion per second, not query latency, not the CPU cost of merges. Disk only. The merge cost of full-text indexes is a topic with a life of its own and is left for the saturation runbook.&lt;/p>
&lt;p>And it does not measure object storage, which in version 4 is mandatory and keeps a copy of the original event. That is additional disk this benchmark does not touch.&lt;/p>
&lt;h2 id="checklist">Checklist&lt;/h2>
&lt;ul>
&lt;li>Do not size using the official container sizing: it covers CPU and memory, not disk growth.&lt;/li>
&lt;li>Count &lt;strong>2,782 bytes per observation&lt;/strong> as a starting point, and adjust by measuring your own load.&lt;/li>
&lt;li>Multiply by the number of replicas before asking for the volume.&lt;/li>
&lt;li>Multiply what the columns suggest by 1.7, because the indexes are 41.8 % of the full table&amp;rsquo;s disk.&lt;/li>
&lt;li>Measure your own installation with the system tables instead of trusting this article: active parts for the total, columns for the breakdown, skip indexes for what is not data.&lt;/li>
&lt;li>Watch the metadata, which is the only thing the materialized view&amp;rsquo;s truncation does not control, because it trims each element and not the whole.&lt;/li>
&lt;li>Size by turn depth, not by number of requests.&lt;/li>
&lt;li>Switch on both crons, the retention one and the deleted mask one, and check that the disk really goes down after the second.&lt;/li>
&lt;li>Reckon on per-project retention requiring a licence, and budget the work of writing it if you do not have one.&lt;/li>
&lt;li>Remember that the tables have no expiry in the engine: with no cron, they grow for ever.&lt;/li>
&lt;/ul>
&lt;h2 id="traps-and-things-that-are-not-what-they-look-like">Traps and things that are not what they look like&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>Langfuse publishes no capacity figure at all&lt;/strong>, neither in its code nor in its documentation.&lt;/li>
&lt;li>&lt;strong>A compression ratio without saying what was compressed means nothing&lt;/strong>: the same schema gives 38 times with repeated text and 3.68 with realistic text.&lt;/li>
&lt;li>&lt;strong>41.8 % of the full table&amp;rsquo;s disk is indexes&lt;/strong>, not data.&lt;/li>
&lt;li>&lt;strong>Indexing the input costs almost as much as storing it&lt;/strong>: 523 bytes per observation against 717.&lt;/li>
&lt;li>&lt;strong>The 78-column skeleton is irrelevant&lt;/strong>: seventeen bytes per observation across the sixty-seven small columns.&lt;/li>
&lt;li>&lt;strong>In the listings table metadata costs more than input and output together.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>The materialized view&amp;rsquo;s truncation trims each element of the metadata, not the whole array.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>The listings table is not an appendix&lt;/strong>: it adds 18.5 % of the first table&amp;rsquo;s disk despite storing two hundred characters.&lt;/li>
&lt;li>&lt;strong>A delete frees no disk.&lt;/strong> The mask has to be applied, and that cron ships disabled.&lt;/li>
&lt;li>&lt;strong>The retention one also ships disabled&lt;/strong> out of the box.&lt;/li>
&lt;li>&lt;strong>The cleaner works table by table&lt;/strong>; the materialized view does not propagate deletes.&lt;/li>
&lt;li>&lt;strong>The event tables carry no expiry clause.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Per-project retention is a paid feature&lt;/strong> when self-hosted.&lt;/li>
&lt;li>&lt;strong>Generations are 11 % of the observations and 65 % of the bytes.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Bloom filters are free&lt;/strong> as far as disk goes; the full-text ones are not.&lt;/li>
&lt;/ul>
&lt;h2 id="closing">Closing&lt;/h2>
&lt;p>What I take from this benchmark is not the figure, although the figure was needed. It is that the split is not where you look for it.&lt;/p>
&lt;p>When somebody sizes a trace database, they look at the size of the prompts. It is the natural thing to do: they are the large fields and they are the ones you see in the interface. And it turns out that in this architecture the prompts are half the compressed data, and the compressed data is a little over half the disk. The rest is search structures nobody counts because they do not show up when you look at a row.&lt;/p>
&lt;p>There is a pattern there that goes beyond this tool. An observability database is, by definition, a database queried in unforeseen ways: by free text, by user, by session, by whatever is needed when something goes wrong at three in the morning. That ability to search by anything is paid for in disk, always, and it is paid at write time and not at search time. Langfuse&amp;rsquo;s design pays it fairly explicitly, with four full-text indexes declared in the migration, in plain sight for anyone who wants to read them.&lt;/p>
&lt;p>What was missing is for somebody to put the number alongside. Now it is there, with its margin of error and its method on the table, so that the next person to measure it can tell me where I got it wrong.&lt;/p>
&lt;h2 id="the-series-the-ten-articles">The series: the ten articles&lt;/h2>
&lt;ol>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/langfuse-v4-what-goes-into-a-trace/">What goes into a trace&lt;/a>: version 4&amp;rsquo;s data model, limits, precedences, scores, masking and indexes.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/langfuse-v4-instrumenting-langgraph/">Putting LangGraph in front&lt;/a>: instrumenting an agentic platform and the measured cost of one turn.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/langfuse-v4-real-agents-odr-gptr/">Two well-known agents instrumented&lt;/a>: Open Deep Research and GPT Researcher, with the measured figures of a real request.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/langfuse-v4-migrating-with-no-window/">Migrating from version 3 to 4 with no window&lt;/a>: the three write-mode steps, the resumable background migrations and where the rollback point of no return sits.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/langfuse-v4-worker-queues/">The worker queues&lt;/a>: the map of all thirty-nine, which pool to dedicate to each group, the per-queue switches, sharding and concurrency.&lt;/li>
&lt;li>Real ClickHouse capacity and cost (this article): how to measure bytes per observation with the system tables, the difference between the full table and the listings table, and the merge cost of full-text indexes.&lt;/li>
&lt;li>Retention, deletion and data protection: why a deletion does not free disk, the mask cleaner that ships disabled, the pending deletion queue and the S3 lifecycle that has to be implemented by hand.&lt;/li>
&lt;li>Backup and cross recovery: restore ordering across Postgres, ClickHouse and object storage, what each mismatch breaks, and how far event replay goes.&lt;/li>
&lt;li>Saturation runbook: what to alert on from the queue metrics, the stuck probes, draining through the readiness endpoint and the dead letter queue.&lt;/li>
&lt;li>Getting the data out: the blob storage integration to Parquet, batch exports and the metrics API, to build the data lake.&lt;/li>
&lt;/ol>
&lt;h2 id="see-also">See also&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/langfuse-inside-sorting-centre-bottleneck/">Self-hosted Langfuse: architecture and tuning&lt;/a> — the architecture of the previous version, where this setup comes from.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/litellm-langfuse-operational-pair/">LiteLLM and Langfuse: the operational pair&lt;/a> — who produces the traces measured here.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/sizing-agents-decode-bottleneck/">Sizing for agents&lt;/a> — the other side of sizing, that of the gateway and the inference engine.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/deepagents-free-sdk-licensed-server/">deepagents on a cluster of your own&lt;/a> — the harness that generates these observations and where it keeps its own state.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/technical-controls-ens-iso-42001-eu-ai-act-cross-mapping/">Technical controls for ENS, 42001 and the AI Act&lt;/a> — the framework where retention and deletion fit.&lt;/li>
&lt;/ul>
&lt;h2 id="sources">Sources&lt;/h2>
&lt;ul>
&lt;li>Langfuse, migraciones de ClickHouse (&lt;code>packages/shared/clickhouse/migrations/canonical/&lt;/code>, ficheros 0039 a 0047): &lt;a href="https://github.com/langfuse/langfuse">https://github.com/langfuse/langfuse&lt;/a>.&lt;/li>
&lt;li>Langfuse, escritura del worker en ClickHouse (&lt;code>worker/src/services/IngestionService/&lt;/code>, &lt;code>worker/src/services/ClickhouseWriter/&lt;/code>): &lt;a href="https://github.com/langfuse/langfuse">https://github.com/langfuse/langfuse&lt;/a>.&lt;/li>
&lt;li>Langfuse, limpiador de retención por lotes y limpiador de máscaras de borrado (&lt;code>worker/src/features/batch-data-retention-cleaner/&lt;/code>, &lt;code>worker/src/features/deleted-mask-cleaner/&lt;/code>): &lt;a href="https://github.com/langfuse/langfuse">https://github.com/langfuse/langfuse&lt;/a>.&lt;/li>
&lt;li>Langfuse, guía de autoalojamiento y requisitos de infraestructura: &lt;a href="https://langfuse.com/self-hosting">https://langfuse.com/self-hosting&lt;/a>.&lt;/li>
&lt;li>ClickHouse, tablas del sistema &lt;code>parts&lt;/code>, &lt;code>columns&lt;/code> y &lt;code>data_skipping_indices&lt;/code>: &lt;a href="https://clickhouse.com/docs/en/operations/system-tables">https://clickhouse.com/docs/en/operations/system-tables&lt;/a>.&lt;/li>
&lt;li>ClickHouse, índices de salto y de texto completo: &lt;a href="https://clickhouse.com/docs/en/engines/table-engines/mergetree-family/mergetree">https://clickhouse.com/docs/en/engines/table-engines/mergetree-family/mergetree&lt;/a>.&lt;/li>
&lt;li>ClickHouse, borrados ligeros y máscara de borrado: &lt;a href="https://clickhouse.com/docs/en/guides/developer/lightweight-delete">https://clickhouse.com/docs/en/guides/developer/lightweight-delete&lt;/a>.&lt;/li>
&lt;/ul></description></item><item><title>deepagents on a cluster of your own: the SDK is free, the server is not, and what you have to build in between</title><link>https://blog.lo0.es/en/posts/deepagents-free-sdk-licensed-server/</link><pubDate>Tue, 15 Sep 2026 07:00:00 +0200</pubDate><guid>https://blog.lo0.es/en/posts/deepagents-free-sdk-licensed-server/</guid><description>&lt;blockquote>
&lt;p>Opening of the agentic platform vertical. The previous articles dealt with the pieces separately: &lt;a href="https://blog.lo0.es/en/posts/choosing-oss-gateway-llm-inference/">the inference gateway&lt;/a>, &lt;a href="https://blog.lo0.es/en/posts/litellm-mcp-gateway-second-front-door/">its second door&lt;/a>, &lt;a href="https://blog.lo0.es/en/posts/sizing-agents-decode-bottleneck/">what a fleet of agents costs&lt;/a> and &lt;a href="https://blog.lo0.es/en/posts/isolating-ai-agents-workstation-to-cluster/">isolation&lt;/a>. This one deals with what goes on top. Verified against &lt;code>deepagents&lt;/code> 0.7.14 and the LangGraph repository, commits of 14 September 2026.&lt;/p>
&lt;/blockquote>
&lt;h2 id="tldr">TL;DR&lt;/h2>
&lt;p>&lt;strong>deepagents is not a runtime.&lt;/strong> It is a factory that assembles middleware and delegates to LangChain&amp;rsquo;s agent factory. It returns a compiled LangGraph graph and nothing else. The repository&amp;rsquo;s own documentation says so without decoration: it introduces no new runtime. That is good news for whoever deploys it, because it runs wherever Python runs.&lt;/p>
&lt;p>&lt;strong>The planning tool no longer exists.&lt;/strong> It disappeared in 0.7.0 as a breaking change, along with the state channel that backed it. Almost everything written about this harness still describes it as one of its three legs. Today you have to ask for it explicitly, and you have to ask for it in each subagent too.&lt;/p>
&lt;p>&lt;strong>With the default configuration, the agent&amp;rsquo;s files live inside the checkpoint.&lt;/strong> The state backend stores the full content, binaries in base64, in the same place the conversation is persisted. Changing that is one line and it is the first architecture decision of the deployment.&lt;/p>
&lt;p>&lt;strong>The database grows per loop step, not per turn.&lt;/strong> Four tables, one row per superstep, one per channel and version, another per task write. And of the four cleanup operations the base class declares, the Postgres implementation has only one: delete an entire thread. The other three raise not-implemented.&lt;/p>
&lt;p>&lt;strong>Pruning by hand can empty the history silently.&lt;/strong> The code itself warns that a partial delete breaks the delta chain and leaves the channels rebuilding empty without raising any error. The only safe pruning is by whole thread.&lt;/p>
&lt;p>&lt;strong>Tenant isolation is the thread identifier, and there is no owner check.&lt;/strong> Whoever can send a thread identifier reads that thread. The on-disk filesystem backend fixes its root when the agent is built, so it separates nobody. And the documentation&amp;rsquo;s example for computing the per-user namespace does not work outside the managed platform.&lt;/p>
&lt;p>&lt;strong>There is a configuration trap that leaves a short-window model uncompacted.&lt;/strong> If the model does not expose its window, the compaction threshold falls back to 170,000 tokens. Served on a vLLM with a 32,768 window, the agent blows up on context length long before it compacts. The failure is silent and is fixed by passing the profile by hand.&lt;/p>
&lt;p>&lt;strong>Observability without the platform does exist, and it goes through a package that is not obvious.&lt;/strong> LangChain&amp;rsquo;s core does not emit OpenTelemetry: zero references. The exporter lives inside the LangSmith SDK, and there is a mode that emits over OTLP only, without talking to the service. It is the road to a Langfuse of your own, and it is worth knowing where it passes.&lt;/p>
&lt;p>&lt;strong>And here is where the free part ends: the server.&lt;/strong> The SDK, the graph engine and the checkpointers are MIT. The server that publishes them over HTTP carries an Elastic 2.0 licence, asks for a commercial licence key for production, adds Redis to the stack and reports metadata to an external endpoint unless there is an isolation agreement. The development tool is explicitly for development.&lt;/p>
&lt;p>&lt;strong>What is left is a list of things you have to write.&lt;/strong> Per-thread authorisation, retention, a run queue, cancellation, clean draining and migrations. None of them is hard. All of them are work, and none of them appears in the twenty-line front-page example.&lt;/p>
&lt;h2 id="you-are-here-the-harness-arrives-after-the-model">You are here: the harness arrives after the model&lt;/h2>
&lt;p>The sequence repeats on every project. First you set up inference, which is the visible part. Then the gateway, because you need keys and budgets. Then observability, because nobody knows what is going on. And when all of that works, somebody asks why the assistant cannot do multi-step tasks, and the word agent enters the scene.&lt;/p>
&lt;p>At that point the temptation is to write the loop by hand, and for two weeks it looks like a good idea. The loop is easy. What is not easy is everything you discover afterwards: what happens when the context fills up, where the intermediate files are stored, how you resume a task that was left half done, how you ask a human for confirmation without losing state.&lt;/p>
&lt;p>An agent harness is exactly that: the packaged answer to those questions. And choosing one looks like a library decision, about the size of choosing an HTTP client. It is not. It decides where your users&amp;rsquo; state lives, which database grows, what can be isolated and what cannot, and in the case at hand, whether there is a commercial licence waiting at the end of the road.&lt;/p>
&lt;p>This article looks at &lt;code>deepagents&lt;/code> with those eyes. It is not a usage guide, and I do not assess whether its agents are any good. I look at what it forces you to build.&lt;/p>
&lt;h2 id="the-analogy-the-scaffolding-and-the-crane">The analogy: the scaffolding and the crane&lt;/h2>
&lt;p>A building site needs scaffolding and it needs a crane, and they are not the same thing even though both are metal structures surrounding the building.&lt;/p>
&lt;p>You put up the scaffolding yourself with parts you bought. It adapts to the façade, it can be extended, and the day the work finishes you take it down and put it away. Nobody charges you for using it once it is bought.&lt;/p>
&lt;p>The crane is another matter. You hire it, it comes with an operator, it has a serial number and you have to notify somebody when it goes up. Without a crane the work still moves forward, more slowly and with more people, but it moves forward.&lt;/p>
&lt;p>Here the scaffolding is the SDK and the graph engine, and it belongs to whoever buys it. The crane is the server that exposes the agent over HTTP, and that one comes with a contract. This article is about where exactly the line between the two sits, because you cannot see it in the brochure, and about how much work can be done without a crane.&lt;/p>
&lt;h2 id="what-deepagents-actually-is">What deepagents actually is&lt;/h2>
&lt;p>The first thing to get out of your head is that it is a system. It is a function.&lt;/p>
&lt;p>The main factory assembles a list of middleware and calls LangChain&amp;rsquo;s agent factory, which is what builds the graph. What it returns is a compiled graph, with an added configuration on top that raises the recursion limit to a very high number and adds an integration tag in the metadata. The repository&amp;rsquo;s own architecture document states it: no new runtime is introduced.&lt;/p>
&lt;p>Compiled with the default options, the graph has the nodes you would expect from a reasoning-and-action loop: start, a prior node that repairs incomplete tool calls, the model node, the tools node and the end. The conditional edges run from the model to tools, to the model itself or to the end. There is nothing more. Everything that looks like a system is middleware that adds nodes only when it declares hooks of its own.&lt;/p>
&lt;p>The tools that ship by default are the filesystem ones, plus the subagent delegation one: list, read, write, edit, delete, search by pattern, search by content, execute, and the task tool.&lt;/p>
&lt;p>The good consequence of this is portability. If something is a LangGraph graph and nothing else, it runs in any Python container, with the checkpointer you choose, against the model you choose. There is no hidden service behind it.&lt;/p>
&lt;h3 id="what-is-no-longer-there">What is no longer there&lt;/h3>
&lt;p>It deserves its own section because it contradicts almost everything published. The planning tool, the one that wrote a task list and kept it in a state channel, &lt;strong>was removed in 0.7.0&lt;/strong> as a breaking change declared in the changelog. The tool is gone, the channel is gone, and the prompt fragment that went with it is gone.&lt;/p>
&lt;p>It survives as LangChain middleware, and you have to ask for it by hand. With one detail that bites: you have to ask for it in each subagent too, because each subagent assembles its own stack. All that remains from the previous stage are archaeological leftovers, such as the state key that is still on the list of keys excluded when passing context to a subagent.&lt;/p>
&lt;p>The plan, today, rests on files. Which is consistent with the rest of the design, and is another reason to look closely at where those files live.&lt;/p>
&lt;h2 id="where-state-lives">Where state lives&lt;/h2>
&lt;p>This is the first architecture decision and the one that is most expensive to get wrong.&lt;/p>
&lt;p>The agent&amp;rsquo;s state has two large channels. Messages, with a reducer of its own that deduplicates by identifier and treats deletions as tombstones. And files, with its own reducer where a null value means deleted. Both use a channel type that writes deltas and stores a full snapshot every fifty updates, and the code comment explains why: to take checkpoint growth from quadratic to linear with respect to the number of messages.&lt;/p>
&lt;p>That channel type is marked beta in its own code, with a warning that the on-disk representation may change. Worth knowing before you rest a platform on top of it.&lt;/p>
&lt;h3 id="the-backends">The backends&lt;/h3>
&lt;p>There is a common interface and several implementations. The ones that matter for deciding:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Backend&lt;/th>
&lt;th>Where the files live&lt;/th>
&lt;th>Persistence&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>State&lt;/td>
&lt;td>Inside the graph state, that is, inside the checkpoint&lt;/td>
&lt;td>Per thread, in the checkpointer&amp;rsquo;s database&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Filesystem&lt;/td>
&lt;td>In a real directory anchored to a root&lt;/td>
&lt;td>Wherever you mount the volume&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Store&lt;/td>
&lt;td>In LangGraph&amp;rsquo;s store, with a namespace&lt;/td>
&lt;td>Crosses threads, with a time to live&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Composite&lt;/td>
&lt;td>Routed by path prefix among the above&lt;/td>
&lt;td>Mixed&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>The default backend is the state one. That is: &lt;strong>by default, everything the agent writes ends up inside the checkpoint&lt;/strong>, and binaries go base64-encoded inside the same JSON. For a demo it is convenient, because you have nothing to set up. For a platform with real users it is a sizing decision taken by omission.&lt;/p>
&lt;p>There is a detail that rounds the matter off. The filesystem middleware automatically evicts large tool results to a file when they exceed a token threshold, so that they do not occupy the context window. It is a good mechanism. But with the state backend, evicting to a file means moving the content from one place in the checkpoint to another place in the same checkpoint. It leaves the model&amp;rsquo;s window and it does not leave the database.&lt;/p>
&lt;h3 id="what-grows-in-the-database">What grows in the database&lt;/h3>
&lt;p>The Postgres checkpointer creates four tables: one for migrations, one for checkpoints, one for blobs by channel and version, and one for writes by task. The primary keys all start with the thread identifier, and there is an index per table on that column.&lt;/p>
&lt;p>What you have to internalise is the rhythm. One checkpoint row is written &lt;strong>per engine superstep&lt;/strong>, not per conversation turn. A turn with four tool calls is several supersteps. None of that is recycled on its own.&lt;/p>
&lt;p>And here comes the uncomfortable part. The checkpointer&amp;rsquo;s base class declares four maintenance operations: delete a thread, delete by run, copy a thread and prune with a strategy. In the Postgres implementation &lt;strong>only the first is implemented&lt;/strong>. The other three inherit the not-implemented exception from the base class. Deleting a thread is three delete statements by identifier, and that is it.&lt;/p>
&lt;p>In other words: retention is yours to build. And with a care that the code itself documents and that is worth quoting in full, because it is the kind of warning you read after the incident. A naive pruning that removes intermediate checkpoints and their writes can cut the delta chain; the surviving checkpoint is rarely a snapshot point, so its channels would rebuild empty &lt;strong>without any error being raised&lt;/strong>.&lt;/p>
&lt;p>Translated into operations: the periodic sweep deletes whole threads and only whole threads. Never loose rows. And since the checkpoints table has no date column, to know which thread is old you have to keep count outside or rely on the checkpoint identifier being a time-sortable UUID.&lt;/p>
&lt;p>There is a variant of the checkpointer that stores only the last state and retains no history, which caps growth at the root. With one important reservation to verify before using it: it is not clear that it is compatible with the delta channels this harness uses by default for messages and files. Combining the two without checking is asking for a silent problem.&lt;/p>
&lt;h3 id="the-store-which-does-have-cleanup">The store, which does have cleanup&lt;/h3>
&lt;p>Curiously, the sibling piece is better resolved. The store has a configurable time to live, with refresh on read, skipping of expired entries and a sweep interval, and the sweep is a delete by expiry date.&lt;/p>
&lt;p>Two warnings. The sweeper is a thread inside the process, so with several replicas there are several sweepers competing for the same delete; it is idempotent, but it is free contention and you avoid it by switching it off and putting the delete in a scheduled cluster job. And the store&amp;rsquo;s semantic search requires the vector extension in the database, which constrains the managed Postgres image, which normally does not carry it.&lt;/p>
&lt;h2 id="the-isolation-that-does-not-exist">The isolation that does not exist&lt;/h2>
&lt;p>If the platform is going to have more than one tenant, this is the section that decides the deployment topology.&lt;/p>
&lt;p>The only thing separating one thread from another is the thread identifier. It is the first column of the primary key in all three tables. &lt;strong>There is no owner check anywhere in the checkpointer.&lt;/strong> Whoever manages to send a thread identifier in the call configuration reads that thread. Authorisation is the responsibility of the layer you write on top, and since that layer also has to be written, it is worth noting down now.&lt;/p>
&lt;p>By backend, the situation differs:&lt;/p>
&lt;ul>
&lt;li>The state one inherits thread isolation, which is enough if the HTTP layer validates that the user owns the thread.&lt;/li>
&lt;li>The store one computes the namespace per call, through a function that receives the runtime context. It is the correct option for multi-tenant.&lt;/li>
&lt;li>The filesystem one fixes its root &lt;strong>in the constructor&lt;/strong>, when the agent is built. There is no path in the code that recomputes it per request, per thread or per tenant. Sharing one deployment between tenants with this backend means everybody sees the same directory.&lt;/li>
&lt;/ul>
&lt;p>On that last point, do not confuse it with the protection that does exist. Virtual mode, which now comes enabled by default, blocks paths that escape the root. That protects against directory traversal. It does not protect against the neighbour. The project&amp;rsquo;s own threat document admits it: that mode exists to support the composite backend&amp;rsquo;s prefix routing, not as a security boundary.&lt;/p>
&lt;p>There is also a documentation trap worth flagging because it costs an afternoon. The example the documentation proposes for computing the per-user namespace reads the identity from a server-information field of the runtime context. That field is annotated in the LangGraph code as metadata injected by the managed server, and &lt;strong>null when running open-source LangGraph without managed deployments&lt;/strong>. On a cluster of your own with an HTTP layer of your own, that example fails. What you have to use is the context you fill in yourself, with the identifier that comes validated from the identity provider.&lt;/p>
&lt;p>And one more boundary that gets crossed without warning: the memory and skills middleware interpolate file content into the system prompt as is, without sanitising. If two tenants share a directory that serves as the source for that, the first writes instructions that the second&amp;rsquo;s agent executes. It is prompt injection through shared storage, and it is avoided with the same measure as everything above: do not share the filesystem backend between tenants.&lt;/p>
&lt;p>The topology conclusion is short. Real multi-tenant with files means &lt;strong>one deployment per tenant&lt;/strong>, with its namespace, its volume and its quota, or else the store backend with a namespace computed per request. There is no third way.&lt;/p>
&lt;h2 id="the-model-and-the-trap-you-need-to-know-about">The model, and the trap you need to know about&lt;/h2>
&lt;p>The default model is a specific commercial provider, and it is marked deprecated: passing an empty model warns and will stop working. You can inject any LangChain chat model, which is what anyone serving their own weights will do.&lt;/p>
&lt;p>With two integration details that are not obvious.&lt;/p>
&lt;p>The first. To point at an OpenAI-compatible endpoint, which is how a vLLM behind a gateway looks, you have to pass &lt;strong>the already-built instance&lt;/strong>, not the string with the provider prefix. With the string, a provider profile is applied that forces the use of the responses API, which a vLLM normally does not implement. Passing the built object applies no profile and the agent stays clean.&lt;/p>
&lt;p>The second is the trap, and in my view it is the most likely configuration failure in the whole deployment.&lt;/p>
&lt;p>The compaction middleware computes its thresholds from the model profile. If the model declares its input window, it compacts on reaching 85 % of the window and keeps 10 %. If it does not declare it, it falls back to a fixed value: &lt;strong>170,000 tokens&lt;/strong>.&lt;/p>
&lt;p>The chain that leads to not declaring it is the usual one in a deployment of your own. The OpenAI client resolves the profile by looking the model name up in a static table of the provider&amp;rsquo;s models. A model served as &lt;code>qwen3-30b-a3b&lt;/code> is not in that table, so the profile comes out empty, and empty becomes null. The method that assigns it swallows any exception, so there is no warning.&lt;/p>
&lt;p>Result: a vLLM with a 32,768-token window will not compact until 170,000. The server will return a context-length error long before that, and it will do so intermittently, only when conversations get long. There is a safety net that catches the error and trims, but that is reactive recovery after a failed call, not planning.&lt;/p>
&lt;p>The fix is one line, because the profile is a public field and only fills itself in when it is empty: you pass it when building the model, with a value below the real window to leave room for generation. And it deserves an &lt;code>assert&lt;/code> at start-up, because it is the kind of thing nobody looks at until it fails.&lt;/p>
&lt;p>A related sizing note. With the default thresholds and a 32,768 window, compaction would trigger at 27,852 tokens, while eviction of large tool results is fixed at 20,000. That is, a single tool result can occupy almost the whole budget before anything evicts it. With 128,000 windows the defaults are proportionate; below 64,000 you have to lower those two thresholds by hand.&lt;/p>
&lt;h2 id="the-tools-and-the-mcp-that-is-not-where-it-seems">The tools, and the MCP that is not where it seems&lt;/h2>
&lt;p>This harness has its filesystem tools and its delegation tool, and the rest are passed as an ordinary list of LangChain tools.&lt;/p>
&lt;p>The fact that matters to anyone with an MCP gateway in place: &lt;strong>the SDK does not integrate MCP&lt;/strong>. There is no reference to the protocol in the package, nor any dependency on adapters. The integration lives in the command-line tool that accompanies the project, which does depend on the MCP adapters for LangChain, resolves configurations in the style of the desktop clients, and turns each remote tool into a LangChain tool that then goes through the ordinary list.&lt;/p>
&lt;p>So the path exists and is proven, but you have to walk it yourself: load the gateway&amp;rsquo;s tools with the adapters and pass them in the list. Which, incidentally, fits well with &lt;a href="https://blog.lo0.es/en/posts/mcp-tool-priority-visibility-and-exclusion/">what we already know about the catalogue&lt;/a>: the credential-based filtering the gateway does is applied in the listing, so each agent receives the subset that belongs to it without the harness having to know anything.&lt;/p>
&lt;h3 id="the-subagents">The subagents&lt;/h3>
&lt;p>They are declared in three ways and invoked with a task tool that takes a description and a type. What you need to know to design with them:&lt;/p>
&lt;p>The subagent receives the parent&amp;rsquo;s state &lt;strong>except&lt;/strong> the messages and a few private keys, and it starts with a single human message which is the task description. Files &lt;strong>are&lt;/strong> inherited, and it shares the same backend object, so parent and child work on the same logical filesystem. The result comes back as a tool message with the structured response or the child&amp;rsquo;s last text.&lt;/p>
&lt;p>Depth is effectively one level. Verified by compiling the graph: the stack assembled for a declarative subagent does not include the subagent middleware, so the child does not have the task tool and cannot delegate in turn.&lt;/p>
&lt;p>And there is no concurrency limit of its own. The parallelism is that of the engine&amp;rsquo;s parallel tool calls, and the task tool&amp;rsquo;s prompt explicitly invites launching them in parallel. Anyone with a per-tenant GPU quota will want to put the brake where it belongs, because it is not here.&lt;/p>
&lt;h2 id="code-execution">Code execution&lt;/h2>
&lt;p>The sandbox protocol is well defined and the packaged implementations are nearly all paid services: four commercial providers, plus an embedded JavaScript interpreter and a sandbox from the platform vendor itself.&lt;/p>
&lt;p>That leaves a local shell backend that is worth looking at closely before considering it. It runs the command with the system shell, and its own docstring lists what it does not do: no isolation, no process separation, no resource limits, and it lists production and multi-tenant environments as inappropriate use cases. It also has an option to inherit the process environment which, enabled, hands the model every variable, including the gateway credentials and the database connection string. It comes disabled, and it is an easy foot to shoot.&lt;/p>
&lt;p>The reasonable way out on a cluster of your own is to implement the protocol yourself, and it is smaller than it looks. The sandbox base class requires four things: an identifier, execute, upload files and download files. Everything else, list, read, write, edit, delete and search, is built by the base class on top of execute. Looking at one of the commercial implementations, it is on the order of a hundred and fifty lines.&lt;/p>
&lt;p>One image requirement that constrains the Dockerfile and is worth knowing beforehand: the base class helpers inject small encoded Python scripts to resolve searches and checks, so the sandbox image &lt;strong>needs Python and a POSIX shell&lt;/strong> or half the calls fail.&lt;/p>
&lt;p>On how to implement it, the shape that fits a Kubernetes of your own is an ephemeral container per session with a bounded time to live, resource limits, no service account token mounted and a network policy that only allows egress towards the inference gateway. Execute is resolved against the cluster&amp;rsquo;s exec API, and upload and download through the same channel. The alternative with less code is an isolated runtime class for the agent pod, one pod per tenant, and the shell backend inside; you lose multi-tenancy in a single pod and you have nothing to write.&lt;/p>
&lt;h2 id="observability-and-the-package-it-goes-through">Observability, and the package it goes through&lt;/h2>
&lt;p>Here there is a fact that surprises and that is worth being clear about before designing the trace pipeline.&lt;/p>
&lt;p>&lt;strong>LangChain&amp;rsquo;s core does not emit OpenTelemetry.&lt;/strong> Zero references in the whole package. Its native tracer talks to the vendor&amp;rsquo;s observability service and that is that.&lt;/p>
&lt;p>The OTLP exporter does exist, but it lives inside that service&amp;rsquo;s SDK. And there is the good news: that SDK has a tracing mode that emits &lt;strong>only&lt;/strong> over OTLP, without talking to the service, reading the destination and headers from OpenTelemetry&amp;rsquo;s standard environment variables. The transport is HTTP with protobuf, which is exactly what a self-hosted Langfuse accepts on its ingestion route.&lt;/p>
&lt;p>That is: to take the agent&amp;rsquo;s traces to your own observability you have to install the vendor&amp;rsquo;s package and ask it not to talk to the vendor. It works, it is supported and it is documented, but it is a dependency you have to declare in the supply chain analysis, not a configuration detail. And if the OpenTelemetry packages are missing, it emits a warning and silently stops tracing, so start-up should check for it.&lt;/p>
&lt;p>There is the alternative of the callback handler that Langfuse itself publishes, with fewer pieces. The practical difference is correlation: over the OTLP route the agent&amp;rsquo;s spans share a trace identifier with the gateway&amp;rsquo;s and the inference engine&amp;rsquo;s, and you see a whole request end to end. Over the callback route, you do not. For anyone who already has a collector deployed, the OTLP route is the one that pays.&lt;/p>
&lt;h2 id="the-server-here-is-where-the-free-part-ends">The server: here is where the free part ends&lt;/h2>
&lt;p>Everything above is an architecture decision. This is a licence decision, and it is the one that gives the article its title.&lt;/p>
&lt;p>The harness SDK is MIT. The graph engine is MIT. The memory, SQLite and Postgres checkpointers are MIT. All of that runs anywhere, with no key, no registration and without calling anybody.&lt;/p>
&lt;p>The server that exposes a graph over HTTP is another matter. The package that implements it carries an &lt;strong>Elastic 2.0 licence&lt;/strong>, whose relevant clauses forbid offering the software to third parties as a managed service and forbid modifying or circumventing the licence key functionality. The command-line tool&amp;rsquo;s own command prints the conditions: for local development it asks for a service API key, and &lt;strong>for production use it asks for a licence key&lt;/strong> in an environment variable.&lt;/p>
&lt;p>There is also usage reporting. The server code contains a beacon endpoint to which it sends metadata, with a header carrying the licence key. The self-hosting documentation confirms it from the other side, listing among the requirements network egress to that domain for licence verification and usage reporting &lt;strong>if not running in isolated mode&lt;/strong>. Isolated mode is a contractual variant, not a box you tick.&lt;/p>
&lt;p>And it drags infrastructure along: besides Postgres, it asks for Redis.&lt;/p>
&lt;p>The development tool, which is the one that appears in every tutorial, describes itself as development mode with hot reload and an in-memory server. It is not a production server and it does not claim to be.&lt;/p>
&lt;p>None of this is a reproach. It is a legitimate and fairly common business model: permissive core, commercial control plane. What is not legitimate is finding out late. Anyone building a platform with a sovereignty requirement, with no dependencies on external services and with the argument that the whole stack is open, has to know that &lt;strong>the piece that turns the graph into a service does not meet that requirement&lt;/strong>.&lt;/p>
&lt;h2 id="what-you-have-to-build">What you have to build&lt;/h2>
&lt;p>The way out is to write the HTTP layer yourself, and the good news is that it is boring. A service that compiles the graph once with the checkpointer pointing at the database, and exposes an endpoint that validates the identity provider&amp;rsquo;s token, derives the thread identifier and the tenant context, and invokes the graph. Event streaming if needed.&lt;/p>
&lt;p>What you additionally have to write, and which the commercial server gave you ready-made:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Per-thread authorisation.&lt;/strong> The engine does not have it. It is the most important and the easiest to forget, because nothing fails if it is missing.&lt;/li>
&lt;li>&lt;strong>Retention and deletion.&lt;/strong> A scheduled job that deletes whole threads, never loose rows, for the reasons in the state section.&lt;/li>
&lt;li>&lt;strong>Migrations.&lt;/strong> Schema setup has to be called explicitly and needs data-definition permissions, so it goes in a separate job with a different role from the runtime one. And since it creates indexes concurrently, it cannot go through a connection pooler in transaction mode; point it at the direct write service.&lt;/li>
&lt;li>&lt;strong>Health probes and draining.&lt;/strong> The engine has a cooperative draining mechanism, but wiring it to the container&amp;rsquo;s termination signal is your own work.&lt;/li>
&lt;li>&lt;strong>Run queue and cancellation&lt;/strong>, if you need long background runs. This is where the absence is most felt, and the reasonable answer is a queue on the same database rather than another system.&lt;/li>
&lt;li>&lt;strong>Durability decided deliberately.&lt;/strong> There are three modes: persist before the next step, persist in parallel, or persist only on exit. The last one is no good on Kubernetes: if the pod dies, the whole turn is lost.&lt;/li>
&lt;/ul>
&lt;p>On that last point it is worth being explicit, because it affects tool design. The engine stores each task&amp;rsquo;s writes as soon as it finishes, and on resume it reapplies those that were already there and does not repeat them. A tool that finished is not run again. A tool &lt;strong>that was in flight when the pod died&lt;/strong> is, because its write never landed. The semantics are at-least-once, so every tool with an external effect has to be idempotent or carry an idempotency key. There is no exactly-once.&lt;/p>
&lt;p>And the same applies to interrupts for human confirmation: on resume, the node is &lt;strong>re-executed in full&lt;/strong>, so any side effect that sat before the interrupt inside the same node is repeated. The rule is that the interrupt goes first.&lt;/p>
&lt;h2 id="how-it-fits-with-activity-logging">How it fits with activity logging&lt;/h2>
&lt;p>Two things about this setup touch compliance head on, and both come out of earlier sections.&lt;/p>
&lt;p>The first is that the checkpoint is, de facto, a repository of personal data. It contains the whole conversation and, with the default backend, the files the agent has written. If the platform serves identifiable people, that database needs encryption, retention of its own and separate access control, exactly like the gateway&amp;rsquo;s spend table. The difference is that here there is no redaction switch that helps: the state is the state.&lt;/p>
&lt;p>The second is traceability. An agent that delegates to subagents and executes tools produces a chain of actions that you have to be able to reconstruct, and the only place where that chain is complete is the trace. Traces going out over OTLP to your own observability stops being a technical preference and becomes the logging mechanism. It is worth designing it that way from the start, and not adding it afterwards.&lt;/p>
&lt;p>As always in this series, the annex&amp;rsquo;s specific codes have to be checked against the current text before taking them into a compliance document; the approach is developed in &lt;a href="https://blog.lo0.es/en/posts/technical-controls-ens-iso-42001-eu-ai-act-cross-mapping/">the technical controls article&lt;/a>.&lt;/p>
&lt;h2 id="checklist">Checklist&lt;/h2>
&lt;ul>
&lt;li>Decide the filesystem backend before the first demo, because the default one puts the content in the database.&lt;/li>
&lt;li>Pass the model profile by hand with the real window, and check it at start-up.&lt;/li>
&lt;li>Lower the eviction thresholds if the model&amp;rsquo;s window is below 64,000 tokens.&lt;/li>
&lt;li>Build the model as an instance, never as a string with a provider prefix, if there is an OpenAI-compatible endpoint behind it.&lt;/li>
&lt;li>Write per-thread authorisation before opening the service to more than one user.&lt;/li>
&lt;li>One deployment per tenant if you use the filesystem backend; a namespace computed per request if you use the store one.&lt;/li>
&lt;li>Compute that namespace from your own context, not from the server-information field, which is null outside the managed platform.&lt;/li>
&lt;li>A scheduled retention job that deletes whole threads and only whole threads.&lt;/li>
&lt;li>Schema migrations in a separate job, with a role of their own and against the direct write service.&lt;/li>
&lt;li>Synchronous durability, and tools with external effects made idempotent.&lt;/li>
&lt;li>Traces over OTLP with the mode that does not talk to the vendor&amp;rsquo;s service, and a start-up check that the packages are there.&lt;/li>
&lt;li>Put a concurrency limit on subagents if there is a GPU quota involved.&lt;/li>
&lt;li>If you need code execution, implement the sandbox protocol against an ephemeral container; the local shell backend is advised against by its own documentation.&lt;/li>
&lt;li>Note in the supply chain analysis the hard dependencies on commercial providers that the package drags in even when unused.&lt;/li>
&lt;/ul>
&lt;h2 id="traps-and-things-that-are-not-what-they-look-like">Traps and things that are not what they look like&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>The planning tool was removed in 0.7.0.&lt;/strong> You have to add it by hand, and in each subagent too.&lt;/li>
&lt;li>&lt;strong>The default backend stores files inside the checkpoint&lt;/strong>, binaries included in base64.&lt;/li>
&lt;li>&lt;strong>Evicting a large result to a file does not get it out of the database&lt;/strong> with that backend, only out of the context window.&lt;/li>
&lt;li>&lt;strong>One checkpoint is written per superstep&lt;/strong>, not per turn.&lt;/li>
&lt;li>&lt;strong>Of the four declared cleanup operations, only delete-a-thread exists in Postgres.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>A partial pruning can empty the history without raising any error.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>The delta channel is in beta&lt;/strong> and its on-disk representation may change.&lt;/li>
&lt;li>&lt;strong>There is no owner check in the checkpointer.&lt;/strong> The thread identifier is not a credential.&lt;/li>
&lt;li>&lt;strong>The filesystem backend&amp;rsquo;s root is fixed when the agent is built&lt;/strong> and it separates no tenants.&lt;/li>
&lt;li>&lt;strong>The documentation&amp;rsquo;s example for the per-user namespace does not work outside the managed platform.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>The filesystem middleware and the subagent middleware cannot be removed.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>If the model does not declare its window, compaction does not trigger until 170,000 tokens&lt;/strong>, silently.&lt;/li>
&lt;li>&lt;strong>Passing the model as a string with a provider prefix forces an API that a vLLM does not implement.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>The SDK does not integrate MCP&lt;/strong>; the integration lives in the command-line tool.&lt;/li>
&lt;li>&lt;strong>Subagents do not delegate&lt;/strong>: depth is one level.&lt;/li>
&lt;li>&lt;strong>LangChain&amp;rsquo;s core does not emit OpenTelemetry&lt;/strong>; the exporter lives in the vendor&amp;rsquo;s service SDK.&lt;/li>
&lt;li>&lt;strong>The HTTP server carries an Elastic 2.0 licence&lt;/strong>, asks for a licence key for production and reports usage unless there is an isolation agreement.&lt;/li>
&lt;li>&lt;strong>The development tool is for development&lt;/strong>, with an in-memory server.&lt;/li>
&lt;li>&lt;strong>Resumption is at-least-once&lt;/strong> for in-flight tools.&lt;/li>
&lt;/ul>
&lt;h2 id="closing">Closing&lt;/h2>
&lt;p>The question this article started with was whether &lt;code>deepagents&lt;/code> fits into a platform of your own. The answer is that it does, and that it fits better than I expected, precisely because it is less than it appears. A middleware stack on top of a graph is easy to host, easy to understand and easy to replace. What does not fit is the piece next to it.&lt;/p>
&lt;p>That asymmetry is the pattern I have been seeing all year across the ecosystem, and it deserves naming. The core is published under a permissive licence because the core is where you compete for adoption. The control plane is published under a restrictive licence because the control plane is where you charge. It happened with the gateways, it is happening with observability, and it happens here. Anyone building a sovereign platform cannot assess a whole project by the licence of its main repository: you have to look package by package, and look at which of them is the one you are going to need the day the thing has users.&lt;/p>
&lt;p>The good part is that the bill, in this case, is paid in work and not in money. What you have to build to do without the commercial server is an HTTP service with no mystery to it, a retention job, a migration job and an authorisation layer. None of that is research. All of it is a week of an engineer who knows what they are doing, and in exchange you are left with a system that deploys the same way in a school as in a classified datacenter, with no internet egress and with nobody counting runs on the other side.&lt;/p>
&lt;p>What I do recommend is making that list before the first demo and not after. Because the front-page example works in twenty lines, and the twenty lines have the default backend, the model without a profile and no authorisation at all.&lt;/p>
&lt;h2 id="see-also">See also&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/choosing-oss-gateway-llm-inference/">Choosing an inference gateway&lt;/a> — the piece underneath, and the same licence pattern seen somewhere else.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/litellm-mcp-gateway-second-front-door/">LiteLLM&amp;rsquo;s MCP gateway&lt;/a> — where the tools this harness consumes come from.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/mcp-tool-priority-visibility-and-exclusion/">MCP tool priority&lt;/a> — how it is decided which subset of the catalogue each agent sees.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/sizing-agents-decode-bottleneck/">Sizing for agents&lt;/a> — the real load a tool loop generates on the gateway and the engine.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/isolating-ai-agents-workstation-to-cluster/">The contractor with the master key&lt;/a> — the network isolation this article takes as necessary.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/second-cost-vector-ai-agents-durable-execution-temporal/">The second cost vector of agents&lt;/a> — durable execution and what a loop that fails halfway costs.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/technical-controls-ens-iso-42001-eu-ai-act-cross-mapping/">Technical controls for ENS, 42001 and the AI Act&lt;/a> — the framework where logging and retention fit.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/litellm-langfuse-operational-pair/">LiteLLM and Langfuse: the operational pair&lt;/a> — the destination of the OTLP traces the article talks about.&lt;/li>
&lt;/ul>
&lt;h2 id="sources">Sources&lt;/h2>
&lt;ul>
&lt;li>&lt;code>langchain-ai/deepagents&lt;/code>, código y documento de arquitectura: &lt;a href="https://github.com/langchain-ai/deepagents">https://github.com/langchain-ai/deepagents&lt;/a>.&lt;/li>
&lt;li>&lt;code>langchain-ai/deepagents&lt;/code>, registro de cambios de la rama del SDK (eliminación de la herramienta de planificación en la 0.7.0): &lt;a href="https://github.com/langchain-ai/deepagents/blob/main/libs/deepagents/CHANGELOG.md">https://github.com/langchain-ai/deepagents/blob/main/libs/deepagents/CHANGELOG.md&lt;/a>.&lt;/li>
&lt;li>&lt;code>langchain-ai/langgraph&lt;/code>, motor, canales y checkpointers: &lt;a href="https://github.com/langchain-ai/langgraph">https://github.com/langchain-ai/langgraph&lt;/a>.&lt;/li>
&lt;li>&lt;code>langgraph-checkpoint-postgres&lt;/code>, esquema de tablas y operaciones de mantenimiento: &lt;a href="https://github.com/langchain-ai/langgraph/tree/main/libs/checkpoint-postgres">https://github.com/langchain-ai/langgraph/tree/main/libs/checkpoint-postgres&lt;/a>.&lt;/li>
&lt;li>&lt;code>langgraph-checkpoint&lt;/code>, clase base del checkpointer y aviso sobre poda con canales de deltas: &lt;a href="https://github.com/langchain-ai/langgraph/tree/main/libs/checkpoint">https://github.com/langchain-ai/langgraph/tree/main/libs/checkpoint&lt;/a>.&lt;/li>
&lt;li>LangChain, despliegue de servidor autónomo (requisitos de licencia, Postgres, Redis y salida de red): &lt;a href="https://docs.langchain.com/langsmith/deploy-standalone-server">https://docs.langchain.com/langsmith/deploy-standalone-server&lt;/a>.&lt;/li>
&lt;li>LangChain, trazado con OpenTelemetry: &lt;a href="https://docs.langchain.com/langsmith/trace-with-opentelemetry">https://docs.langchain.com/langsmith/trace-with-opentelemetry&lt;/a>.&lt;/li>
&lt;li>Elastic License 2.0: &lt;a href="https://www.elastic.co/licensing/elastic-license">https://www.elastic.co/licensing/elastic-license&lt;/a>.&lt;/li>
&lt;li>Langfuse, integración nativa OpenTelemetry (ruta de ingesta y autenticación): &lt;a href="https://langfuse.com/integrations/native/opentelemetry">https://langfuse.com/integrations/native/opentelemetry&lt;/a>.&lt;/li>
&lt;li>Langfuse, integración con LangChain por retrollamadas: &lt;a href="https://langfuse.com/integrations/frameworks/langchain">https://langfuse.com/integrations/frameworks/langchain&lt;/a>.&lt;/li>
&lt;li>&lt;code>langchain-ai/helm&lt;/code>, chart del servidor: &lt;a href="https://github.com/langchain-ai/helm">https://github.com/langchain-ai/helm&lt;/a>.&lt;/li>
&lt;li>Boletín Oficial del Estado, Real Decreto 311/2022, Esquema Nacional de Seguridad: &lt;a href="https://www.boe.es/buscar/act.php?id=BOE-A-2022-7191">https://www.boe.es/buscar/act.php?id=BOE-A-2022-7191&lt;/a>.&lt;/li>
&lt;/ul></description></item><item><title>Making the model prefer our tools: the priority that exists in no layer, the two primitives that do, and the `_meta` that crosses the gateway</title><link>https://blog.lo0.es/en/posts/mcp-tool-priority-visibility-and-exclusion/</link><pubDate>Tue, 15 Sep 2026 03:30:00 +0200</pubDate><guid>https://blog.lo0.es/en/posts/mcp-tool-priority-visibility-and-exclusion/</guid><description>&lt;blockquote>
&lt;p>A follow-up to &lt;a href="https://blog.lo0.es/en/posts/litellm-mcp-gateway-second-front-door/">LiteLLM&amp;rsquo;s MCP gateway&lt;/a>, which treated the tool catalogue as cost and as surface. This one treats the next question, the one that arrives in the meeting afterwards: so then, how do I get the model to use mine. Verified against LiteLLM 1.102.0 and MCP specification &lt;code>2026-07-28&lt;/code>, with a test of our own run on 14 September 2026.&lt;/p>
&lt;/blockquote>
&lt;h2 id="tldr">TL;DR&lt;/h2>
&lt;p>&lt;strong>Tool priority does not exist as a concept in any layer of the stack.&lt;/strong> Not in the specification, not in the clients, not in the providers&amp;rsquo; APIs. What exist are two distinct primitives: removing the alternative from the catalogue, and leaving it behind a search step. Everything else is persuasion, and the model ignores it when it suits.&lt;/p>
&lt;p>&lt;strong>The specification has no ranking field.&lt;/strong> The &lt;code>Tool&lt;/code> object is name, title, icons, description, schemas, five annotations that are unreliable hints by their own declaration, and &lt;code>_meta&lt;/code>. There is a trap I see repeated: MCP does define a numeric &lt;code>priority&lt;/code>, but in another structure, the one that applies to content blocks, resources and prompts. It does not govern tool selection.&lt;/p>
&lt;p>&lt;strong>The only filtering the specification blesses is by credential.&lt;/strong> The current revision says the set of tools must not vary by connection, and that it may vary according to the authorisation presented. That is exactly what a gateway with per-key lists does.&lt;/p>
&lt;p>&lt;strong>The client is the one that decides, and only one of them has anything resembling priority.&lt;/strong> Claude Code lets you exempt a server from deferral with &lt;code>alwaysLoad&lt;/code>, so that its tools load in full at start-up while the rest require a prior search. It is not called priority and it works as such.&lt;/p>
&lt;p>&lt;strong>We proved that the per-tool mark crosses the gateway.&lt;/strong> We set up an MCP server with one marked tool, served it through LiteLLM 1.102.0 and captured the wire: the &lt;code>_meta&lt;/code> arrives whole and with the correct key. The chain is clean on purpose, with a comment in the code that says so.&lt;/p>
&lt;p>&lt;strong>And it is lost on exactly three routes, which are the gateway&amp;rsquo;s three catalogue-trimming modes.&lt;/strong> Proxy mode, tool search driven by key permissions, and the REST listing. In other words: the per-tool mark and LiteLLM&amp;rsquo;s catalogue trimming are incompatible. You have to choose one.&lt;/p>
&lt;p>&lt;strong>There is a library trap that makes the field disappear silently.&lt;/strong> The SDK&amp;rsquo;s &lt;code>Tool&lt;/code> model declares the alias without allowing population by field name, so building it with the field name puts the dictionary in the extras and serialises it under the wrong key. There is no error. The field simply does not arrive.&lt;/p>
&lt;p>&lt;strong>The cheap lever is the description, not the name.&lt;/strong> Swapping descriptions shifts the selection distribution substantially; changing only the name has minimal and inconsistent effects. And the gateway&amp;rsquo;s overrides really do replace what is sent to the client.&lt;/p>
&lt;p>&lt;strong>Reordering the catalogue is a weak lever.&lt;/strong> The model is already attending to the correct tool eighty per cent of the time when it fails. Interventions on prompt order repair at most 23 % of the failures.&lt;/p>
&lt;p>&lt;strong>And the asymmetry you build is attack surface.&lt;/strong> Whoever controls the order controls the selection: with injection rates of 1.2 % you can take over the head of a tool ranking between 91 % and 97 % of the time.&lt;/p>
&lt;h2 id="you-are-here-the-question-that-arrives-after-the-third-server">You are here: the question that arrives after the third server&lt;/h2>
&lt;p>The previous article closed with a decision ladder for the catalogue: trim, rewrite, split by route, and use progressive disclosure when the catalogue grows. That ladder answers how much catalogue to show.&lt;/p>
&lt;p>The question that arrives afterwards is a different one, and it is more uncomfortable. A real agent has native tools from the client hosting it, it has the in-house gateway with the house servers, and it has two or three third-party servers that somebody connected. All of them compete. And what you want is not to show less catalogue: you want that, faced with a query three tools could serve, yours wins. The audited one, the one with access control in place, the one that leaves a trace.&lt;/p>
&lt;p>The answer everybody gives is to write a rule in the agent&amp;rsquo;s instruction file politely asking it to prefer the in-house server. There is even a rule published in a template marketplace that does exactly that. It is text in the prompt and the model follows it when it feels like it.&lt;/p>
&lt;p>This article is what lies underneath.&lt;/p>
&lt;h2 id="the-analogy-the-counter-with-two-trays">The analogy: the counter with two trays&lt;/h2>
&lt;p>Let us go back to the switchboard operator from the previous article, the one who also hands out keys. Now the building has three key providers: the house cabinet, the maintenance contractor&amp;rsquo;s, and the cleaning service&amp;rsquo;s. All three have a key to the store room and all three open it.&lt;/p>
&lt;p>The head of security wants the house one used, because it is the one with a log. And he discovers he can do three things, not one more.&lt;/p>
&lt;p>He can take the other two off the counter. It always works and it annoys whoever needed them.&lt;/p>
&lt;p>He can leave his own in the front tray and the other two in a drawer, so that to take one you have to ask. It works almost always and annoys nobody.&lt;/p>
&lt;p>And he can change the label on his key so that it describes better when it is the right one. It works sometimes, it is the cheapest, and it is the only one that also protects against the contractor changing his label without warning.&lt;/p>
&lt;p>What he cannot do is put a number from one to ten on each key and trust the operator to respect it. That number does not exist. The rest of the article is why it does not exist and what to do instead.&lt;/p>
&lt;h2 id="what-the-specification-says-nothing">What the specification says: nothing&lt;/h2>
&lt;p>The current revision is &lt;code>2026-07-28&lt;/code>. Read against the source schema, the &lt;code>Tool&lt;/code> object has name, title, icons, description, input schema, output schema, annotations and &lt;code>_meta&lt;/code>. There is no field for priority, weight, rank, cost, group or tag.&lt;/p>
&lt;p>The annotations are exactly five: title, and the read-only, destructive, idempotent and open-world hints. The schema itself warns that all of them are hints and that a client should not make tool-use decisions based on them when they come from untrusted servers.&lt;/p>
&lt;h3 id="the-priority-that-exists-and-is-no-use">The &lt;code>priority&lt;/code> that exists and is no use&lt;/h3>
&lt;p>Here is the mistake I see repeated most. MCP &lt;strong>does&lt;/strong> define a &lt;code>priority&lt;/code> field, numeric between zero and one. It is in the annotations interface that applies to content blocks, to resources and to prompts, alongside the audience and the last-modified date.&lt;/p>
&lt;p>It is not in &lt;code>Tool&lt;/code>. It has no relation whatsoever to which tool the model picks. Anyone who finds it by searching the word in the schema and concludes that tool priority is standardised has the wrong structure.&lt;/p>
&lt;p>The only precedence the specification defines for a tool is one of presentation: to display the name, first the title, then the annotations title, then the name.&lt;/p>
&lt;h3 id="the-only-blessed-filtering-is-by-credential">The only blessed filtering is by credential&lt;/h3>
&lt;p>There is a new sentence in the current revision that is genuinely useful and that goes unnoticed. The set of tools must not vary by connection nor as a side effect of other requests, and it may vary according to the authorisation presented in the request.&lt;/p>
&lt;p>Translated: segmenting the catalogue by key, by team or by token is the legitimate way to do it. Segmenting it by session state stopped being legitimate when the July revision removed protocol sessions.&lt;/p>
&lt;h3 id="discovery-does-not-discover-tools">Discovery does not discover tools&lt;/h3>
&lt;p>&lt;code>server/discover&lt;/code>, which the current revision adds and which servers have to implement, does not return tools. It returns supported versions, capabilities, server instructions, identity and the cache freshness fields. It accepts neither filters nor a cursor.&lt;/p>
&lt;p>Listing is still asking for everything and paginating. The listing request accepts only a cursor: there is no query, no filter, no limit, no groups, no tags.&lt;/p>
&lt;p>And on collisions between servers, the specification explicitly and reasonedly washes its hands: clients or proxies that aggregate tools from several servers will encounter collisions and should implement a disambiguation strategy, for example prefixing. And it adds that the server name is not guaranteed to be unique and should not be used to disambiguate. There is no precedence rule between servers. There is not going to be one soon.&lt;/p>
&lt;h3 id="what-there-is-in-proposals">What there is in proposals&lt;/h3>
&lt;p>Nothing accepted. The official proposals index only lists those in a final state, and none of them deals with priority, groups or search. In draft there is one asking for a free-text query in the listing and a filtering capability; another on groups and tags that ended up superseded; and one carrying the word priority but which is per-tool model routing, which is another thing.&lt;/p>
&lt;p>The interest group exploring the grouping of primitives declares in its own minutes that it is not going to pick a canonical pattern soon, and the documents in its repository are empty. The year&amp;rsquo;s roadmap does not mention the problem.&lt;/p>
&lt;h2 id="the-layer-that-decides-the-client">The layer that decides: the client&lt;/h2>
&lt;p>If the specification gives you nothing, what is left is whatever each client has built on its own. And here there is a clear winner.&lt;/p>
&lt;h3 id="claude-code">Claude Code&lt;/h3>
&lt;p>Tool search is on by default, and what it does is defer: MCP tools are not loaded in full at start-up, the model sees just enough and discovers them when it searches. It requires a model that supports tool reference blocks, that is Claude Sonnet 4.5, Haiku 4.5, Opus 4.5 and later.&lt;/p>
&lt;p>On top of that, the lever:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-json" data-lang="json">&lt;span class="line">&lt;span class="cl">&lt;span class="p">{&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;mcpServers&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;gateway-casa&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;type&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="s2">&amp;#34;http&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;url&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="s2">&amp;#34;https://litellm.interna.svc/mcp&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;alwaysLoad&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="kc">true&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The documentation describes it bluntly: if a server&amp;rsquo;s tools must always be visible without a search step, you set &lt;code>alwaysLoad&lt;/code> to true and all of that server&amp;rsquo;s tools are loaded into context at the start of the session, regardless of the global search setting. And it recommends using it for a small number of tools that the model needs on every turn, because each tool loaded up front consumes context.&lt;/p>
&lt;p>That is visibility asymmetry, and it works as priority even if it is not called that. Ours in front, the rest behind one step.&lt;/p>
&lt;p>There is an operational consequence that the documentation mentions and that matters more in a gateway than in an ordinary server: setting &lt;code>alwaysLoad&lt;/code> makes start-up &lt;strong>wait&lt;/strong> for that server to return its tools, capped by the standard five-second connection timeout. A gateway that aggregates eight upstream servers and does not cache the listing, as is the case, lists live against all eight every time. If one of them is slow, the session&amp;rsquo;s start-up feels it.&lt;/p>
&lt;p>There is also the fine-grained variant. An MCP server can mark individual tools as always loaded by including &lt;code>&amp;quot;anthropic/alwaysLoad&amp;quot;: true&lt;/code> in the tool&amp;rsquo;s &lt;code>_meta&lt;/code> object, with the same effect for that tool alone. That would let you load three tools from our gateway up front and leave the other forty from the same gateway deferred. It is the lever that matters, and it is the one that motivated the test in the next section.&lt;/p>
&lt;p>The rest of this client&amp;rsquo;s arsenal is pure exclusion, and it is worth knowing one detail that decides whether it works or not. A deny rule with the bare tool name &lt;strong>removes it from the model&amp;rsquo;s context&lt;/strong>; with parentheses, it does not. The documentation puts it in a table: plain &lt;code>WebFetch&lt;/code> removes the tool and the model cannot search at all, whereas &lt;code>WebFetch(domain:*)&lt;/code> keeps the tool and rejects every request. The difference is not cosmetic: in the second case the model still sees the tool, still tries it and still spends the context its definition occupies.&lt;/p>
&lt;p>The &lt;code>--tools&lt;/code> flag is the scalpel: it accepts the empty string to disable all the built-ins, and it does not affect MCP tools. That is, &lt;code>--tools &amp;quot;&amp;quot;&lt;/code> leaves the model with only what the gateway serves. It is a command-line flag and has no equivalent as a settings-file key; the functional equivalent there is deny rules.&lt;/p>
&lt;h3 id="the-other-clients">The other clients&lt;/h3>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Client&lt;/th>
&lt;th>Disable the native ones&lt;/th>
&lt;th>Filter the MCP ones&lt;/th>
&lt;th>Priority&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Claude Code&lt;/td>
&lt;td>&lt;code>--tools &amp;quot;&amp;quot;&lt;/code>, deny by bare name&lt;/td>
&lt;td>deny by pattern, subagents&lt;/td>
&lt;td>&lt;code>alwaysLoad&lt;/code>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Codex CLI&lt;/td>
&lt;td>&lt;code>features.shell_tool&lt;/code> to false&lt;/td>
&lt;td>&lt;code>enabled_tools&lt;/code> and &lt;code>disabled_tools&lt;/code> per server&lt;/td>
&lt;td>No&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Gemini CLI&lt;/td>
&lt;td>&lt;code>tools.core&lt;/code> as an allowlist&lt;/td>
&lt;td>&lt;code>includeTools&lt;/code> and &lt;code>excludeTools&lt;/code> per server&lt;/td>
&lt;td>No&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Copilot in VS Code&lt;/td>
&lt;td>Not documented&lt;/td>
&lt;td>picker and tool sets&lt;/td>
&lt;td>No&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Cursor&lt;/td>
&lt;td>No&lt;/td>
&lt;td>per-server switches&lt;/td>
&lt;td>No&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>Two notes. Copilot cuts off at 128 tools per request and its virtualisation threshold does not go above that; its embedding-based routing is internal and has no configuration surface. And Cursor&amp;rsquo;s rules are purely indicative: they are text prepended to the context.&lt;/p>
&lt;h2 id="the-api-layer">The API layer&lt;/h2>
&lt;p>If the client is in-house, the margin is wider, though narrower than it looks.&lt;/p>
&lt;p>&lt;strong>Forcing a set is not possible.&lt;/strong> In Anthropic&amp;rsquo;s case, tool choice accepts auto, any, a specific one by name, or none, and the documentation says explicitly that pointing at a set of MCP tools or at a member is not supported. In vLLM it is the same: a specific one, all of them or none. There is no way to say &amp;ldquo;any of mine from the gateway&amp;rdquo;. The proposal that would introduce grammars restricting tool names to the request&amp;rsquo;s set has been open for months and is not merged.&lt;/p>
&lt;p>&lt;strong>Filtering the catalogue is possible, in all three of the big ones.&lt;/strong> In Anthropic, with an MCP tool-set block where the default configuration disables and specific ones are enabled. In OpenAI, with the allowed tools list inside the MCP block itself, which the documentation justifies on latency and cost grounds, avoiding the model seeing unnecessary definitions. And in Gemini, the remote MCP server configuration accepts restricting which of the server&amp;rsquo;s tools the agent may call.&lt;/p>
&lt;p>It is worth not confusing two things that are named almost identically in OpenAI: the allowed tools list &lt;strong>inside&lt;/strong> the MCP block filters the catalogue, and that is what matters here; the identically named tool-choice type is another thing, it is a selection restriction, and there is evidence that it does not combine well with hosted tools, without the current documentation recording it. If anyone depends on that, let them test it against their model: it fails as a request error, so it shows up on a cold start.&lt;/p>
&lt;p>&lt;strong>And there is a 2026 lever that almost nobody is using.&lt;/strong> Anthropic has in beta a mechanism for mid-conversation tool changes, with add and remove blocks that accept referencing an entire MCP set. The reason it exists is in its documentation: the tools array sits even earlier in the chunked prefix than the system field, so editing it invalidates the cache for the whole conversation; by declaring the full set at the start and using the blocks, the array never changes and the cached prefix stays intact.&lt;/p>
&lt;p>Applied to our case: you can remove a rival server&amp;rsquo;s set hot, mid-session, without paying for the whole prefix. It is not available on every model.&lt;/p>
&lt;h2 id="the-test-the-_meta-crosses-the-gateway">The test: the &lt;code>_meta&lt;/code> crosses the gateway&lt;/h2>
&lt;p>The fine-grained variant of Claude Code&amp;rsquo;s lever depends on a question no documentation answers: whether a gateway that aggregates upstream servers propagates each tool&amp;rsquo;s &lt;code>_meta&lt;/code> all the way to the client, or loses it along the way.&lt;/p>
&lt;p>We checked.&lt;/p>
&lt;h3 id="the-setup">The setup&lt;/h3>
&lt;p>A minimal MCP server with the Python SDK, serving two tools over HTTP, one of them with the mark in place:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-python" data-lang="python">&lt;span class="line">&lt;span class="cl">&lt;span class="n">types&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">Tool&lt;/span>&lt;span class="p">(&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="o">**&lt;/span>&lt;span class="p">{&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="s2">&amp;#34;name&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="s2">&amp;#34;always_loaded_tool&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="s2">&amp;#34;description&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="s2">&amp;#34;tool with _meta&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="s2">&amp;#34;inputSchema&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{&lt;/span>&lt;span class="s2">&amp;#34;type&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="s2">&amp;#34;object&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="s2">&amp;#34;properties&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{}},&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="s2">&amp;#34;_meta&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{&lt;/span>&lt;span class="s2">&amp;#34;anthropic/alwaysLoad&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="kc">True&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="s2">&amp;#34;probe&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="s2">&amp;#34;UPSTREAM_META_MARKER&amp;#34;&lt;/span>&lt;span class="p">},&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">)&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>In front of it, a LiteLLM from the main branch with that server declared in &lt;code>config.yaml&lt;/code>, with no database, with a master key. And an MCP client querying the listing against &lt;code>/mcp&lt;/code>.&lt;/p>
&lt;h3 id="the-result">The result&lt;/h3>
&lt;p>This is what comes out on the wire:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-json" data-lang="json">&lt;span class="line">&lt;span class="cl">&lt;span class="p">{&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;_meta&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{&lt;/span>&lt;span class="nt">&amp;#34;litellm.ai/server_outcomes&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{&lt;/span>&lt;span class="nt">&amp;#34;probe&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{&lt;/span>&lt;span class="nt">&amp;#34;status&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="s2">&amp;#34;ok&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="nt">&amp;#34;tool_count&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="mi">2&lt;/span>&lt;span class="p">}}},&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;tools&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">[&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="p">{&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;name&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="s2">&amp;#34;probe-always_loaded_tool&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;description&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="s2">&amp;#34;tool with _meta&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;inputSchema&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{&lt;/span>&lt;span class="nt">&amp;#34;type&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="s2">&amp;#34;object&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="nt">&amp;#34;properties&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{}},&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;_meta&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{&lt;/span>&lt;span class="nt">&amp;#34;anthropic/alwaysLoad&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="kc">true&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="nt">&amp;#34;probe&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="s2">&amp;#34;UPSTREAM_META_MARKER&amp;#34;&lt;/span>&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="p">]&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The mark arrives whole and with the correct key. The client parses it without trouble.&lt;/p>
&lt;p>Tracing the code, the chain is clean on purpose. LiteLLM keeps the SDK objects as they are when it lists against the upstream, instead of rebuilding them. The name prefixing does a deep copy and mutates only the name, with a comment in the code declaring the intention to preserve every field including the &lt;code>_meta&lt;/code> by avoiding mutation. The two permission filters are list comprehensions that reuse the same objects. And the name and description overrides mutate in place, so they do not lose anything either. Somebody thought about this.&lt;/p>
&lt;p>One operational detail: the name arrives prefixed with the server alias, governed by the configurable separator and the short prefix mode.&lt;/p>
&lt;h3 id="the-three-routes-where-it-is-lost">The three routes where it is lost&lt;/h3>
&lt;p>And here is the finding that changes the previous article&amp;rsquo;s recommendation.&lt;/p>
&lt;p>Proxy mode returns only its three fixed tools, with no &lt;code>_meta&lt;/code> and with no upstream tool at all. Verified against the wire as well.&lt;/p>
&lt;p>Tool search activated by key permissions does the same with four virtual tools.&lt;/p>
&lt;p>And the REST listing builds its response object by hand and discards the field: the tool that did carry a mark comes out with the field null. The class inherits from the type that has the field; it simply is not filled in. The fix fits on one line.&lt;/p>
&lt;p>Those three routes are, exactly, the three ways LiteLLM has of trimming the catalogue. The operational conclusion is uncomfortable and it is worth saying plainly: &lt;strong>either you use the gateway&amp;rsquo;s catalogue trimming, or you use the client&amp;rsquo;s per-tool mark. They cannot be combined.&lt;/strong> The previous article recommended proxy mode as the fifth rung of the ladder; with this in hand, that rung and the fine-grained mark are mutually exclusive.&lt;/p>
&lt;h3 id="the-library-trap">The library trap&lt;/h3>
&lt;p>This deserves its own section because it takes out anyone writing a server or a proxy, and it gives no warning.&lt;/p>
&lt;p>The SDK&amp;rsquo;s tool model declares the field with an alias, and does not enable populating it by name. The consequences, both checked by running them:&lt;/p>
&lt;p>Building with the field name &lt;strong>populates nothing&lt;/strong>. The dictionary slips into the model&amp;rsquo;s extras and comes out serialised under the wrong key, without the leading underscore, which is a key no client interprets.&lt;/p>
&lt;p>And dumping and revalidating without asking for aliases &lt;strong>loses the field&lt;/strong>, for the same reason.&lt;/p>
&lt;p>There is no exception, no warning in the log, nothing. The field disappears. If someone marks their tools and does not see them marked in the client, this is the first place to look, before the gateway.&lt;/p>
&lt;h3 id="what-we-have-not-tested">What we have not tested&lt;/h3>
&lt;p>For honesty&amp;rsquo;s sake, and because it is the link that remains: we have shown that the mark survives the gateway. We have not shown that the client acts on it when the tool arrives with the name prefixed by the gateway. Claude Code&amp;rsquo;s documentation describes the per-tool mark in a single sentence and says nothing about aggregating servers or about prefixes.&lt;/p>
&lt;p>It is a half-hour test for anyone who has the setup in front of them: two tools from the same gateway, one marked, and look at which one appears loaded at the start of the turn. If someone does it before I do, I am interested in the result.&lt;/p>
&lt;h2 id="what-litellm-can-and-cannot-do">What LiteLLM can and cannot do&lt;/h2>
&lt;p>With the above, the gateway&amp;rsquo;s inventory of levers comes out like this.&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Lever&lt;/th>
&lt;th>Serves external MCP clients&lt;/th>
&lt;th>Preserves the &lt;code>_meta&lt;/code>&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>allowed_tools&lt;/code> and &lt;code>disallowed_tools&lt;/code> per server&lt;/td>
&lt;td>Yes&lt;/td>
&lt;td>Yes&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Name and description overrides&lt;/td>
&lt;td>Yes&lt;/td>
&lt;td>Yes&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Tool sets per route&lt;/td>
&lt;td>Yes&lt;/td>
&lt;td>Yes&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Proxy mode and virtual tools&lt;/td>
&lt;td>Yes&lt;/td>
&lt;td>No&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Search by key permissions&lt;/td>
&lt;td>Yes&lt;/td>
&lt;td>No&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>REST listing&lt;/td>
&lt;td>Yes&lt;/td>
&lt;td>No&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Semantic filter&lt;/td>
&lt;td>No&lt;/td>
&lt;td>Not applicable&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>Two clarifications on the previous article&amp;rsquo;s table, now that the code has been read again.&lt;/p>
&lt;p>The allowed and disallowed lists &lt;strong>are applied in the listing&lt;/strong>, not only on the call. They are the blunt, effective instrument, and they do not break the cache because they are static.&lt;/p>
&lt;p>And the overrides have a limit that was not accounted for: they are applied in the MCP protocol listing, but they are &lt;strong>skipped in proxy mode and not applied on the REST route&lt;/strong>. Anyone rewriting descriptions and also turning on proxy mode is not serving what they think.&lt;/p>
&lt;p>What the gateway cannot do is inject &lt;code>_meta&lt;/code>. There is no setting equivalent to the ones for name and description. The mark has to be put there by the origin server, which for in-house servers is trivial because we write them. To mark third-party tools you would have to patch, and it is small: one field in the type and three lines in the overrides function if configuring it by file is enough.&lt;/p>
&lt;p>There is also an ordering lever that did not appear in the previous article. In the proxy&amp;rsquo;s search mode there is a core tools setting that puts them at the front of the ranking and that also does not count against the results cap. It is the closest thing to a declarative priority in the whole stack, and it lives inside the one mode that discards the &lt;code>_meta&lt;/code>.&lt;/p>
&lt;p>And a nuance about proxy mode&amp;rsquo;s results cap. The client cannot ask for more than five, that is true, but the operator can raise it by configuration. The previous article implied it was immovable.&lt;/p>
&lt;h2 id="the-cheap-lever-the-description-not-the-name">The cheap lever: the description, not the name&lt;/h2>
&lt;p>If you cannot exclude, what is left is to bias. And bias comes in through the description.&lt;/p>
&lt;p>The work that measures selection bias between functionally equivalent tools puts the combined bias of the evaluated models between 0.25 and 0.38, that is, you would have to redistribute between 25 % and 38 % of the probability mass for equivalent tools to be picked equally. On that basis, swapping two tools&amp;rsquo; descriptions shifts selection substantially, whereas perturbations that touch only the name produce smaller and more inconsistent effects.&lt;/p>
&lt;p>Another piece of work measures that a single pass of automatic description rewriting improves the metric over a large corpus, almost as much as iterative refinement, and with two orders of magnitude less time. And it adds a warning worth retaining: descriptions tuned against a fixed set of candidates do not generalise to the dynamically retrieved set. You have to optimise for the regime you serve in.&lt;/p>
&lt;p>Applied: the gateway&amp;rsquo;s description override is at once token trimming, selection bias and mitigation of description change after approval. It is the lever with the best effort-to-effect ratio in the whole article, and it is static, so it does not touch the prompt cache.&lt;/p>
&lt;p>Name prefixing, by contrast, costs tokens and probably does not change which tool the model goes to.&lt;/p>
&lt;h2 id="what-does-not-work">What does not work&lt;/h2>
&lt;p>&lt;strong>Reordering the catalogue.&lt;/strong> A June paper measures that, when the model fails, it was already attending to the correct tool eighty per cent of the time, well above chance. The bottleneck is not at the input but in the late layers. Interventions on prompt order recover at most 23 % of the failures, against the 59 % to 91 % of those acting on the final readout. Reordering is cheap and that is why it gets recommended a lot; it is also weak.&lt;/p>
&lt;p>&lt;strong>Rules and instructions.&lt;/strong> They are text. They help and they do not decide.&lt;/p>
&lt;p>&lt;strong>Your own per-request top-K filters.&lt;/strong> This is covered in the previous article and it still holds: four out of five measured strategies perform worse than not filtering, and on top of that rewriting the tool block on every turn invalidates the entire cached prefix.&lt;/p>
&lt;h2 id="the-ladder-updated">The ladder, updated&lt;/h2>
&lt;ol>
&lt;li>&lt;strong>Exclude in the client.&lt;/strong> It is the only deterministic thing. &lt;code>--tools &amp;quot;&amp;quot;&lt;/code> or deny by bare name for the competing native tools.&lt;/li>
&lt;li>&lt;strong>Mark your own as always loaded.&lt;/strong> At server level with &lt;code>alwaysLoad&lt;/code>, or per tool with the mark in the &lt;code>_meta&lt;/code> if you control the origin server. Counting on start-up waiting for the server.&lt;/li>
&lt;li>&lt;strong>Trim with static lists&lt;/strong> per server and per key, and split a catalogue by agent profile across routes.&lt;/li>
&lt;li>&lt;strong>Rewrite the descriptions&lt;/strong> with the gateway&amp;rsquo;s overrides.&lt;/li>
&lt;li>&lt;strong>Remove hot&lt;/strong> the rival sets, if the client is in-house and the model supports it.&lt;/li>
&lt;li>&lt;strong>Choose&lt;/strong>: the gateway&amp;rsquo;s catalogue trimming, or the per-tool mark. Not both.&lt;/li>
&lt;li>&lt;strong>Do not&lt;/strong>: reorder, rename, or trust prompt rules.&lt;/li>
&lt;/ol>
&lt;h2 id="the-risk-that-has-to-be-declared">The risk that has to be declared&lt;/h2>
&lt;p>The asymmetry you build is also a surface. Whoever controls which tools head a ranking controls the selection, and that can be attacked: there is work measuring that by injecting adversarial tools at rates of 1.2 % you take over the head of the ranking between 91 % and 97 % of the time.&lt;/p>
&lt;p>In a system under ENS (Esquema Nacional de Seguridad, Spain&amp;rsquo;s national security framework) this fits under change management, not access control: the set of tools a model sees and the order in which it sees them are security configuration, and today nobody versions them. Neither does the specification fix a hash of the description or of the schema, nor does the gateway. The specific Annex II codes are worth checking against the current text before taking them into a compliance document; the approach is developed in &lt;a href="https://blog.lo0.es/en/posts/litellm-mcp-gateway-second-front-door/">the MCP gateway article&lt;/a>.&lt;/p>
&lt;p>There is a precedent within the ecosystem itself that points the way. The MCP skills extension, which is in a final state, requires a manifest with a per-file hash, mandatory verification before use, and approval tied to the set of files and their digests, so that any change revokes the approval. And it obliges hosts to prevent two skills with the same name from silently replacing one another. It is exactly the pattern missing for tools. Anyone who needs it today has to implement it in their gateway.&lt;/p>
&lt;h2 id="checklist">Checklist&lt;/h2>
&lt;ul>
&lt;li>Decide the strategy before touching anything: catalogue trimming in the gateway, or the per-tool mark in the client. They are mutually exclusive.&lt;/li>
&lt;li>If you pick the mark, put it in the origin server, because the gateway cannot inject it.&lt;/li>
&lt;li>Build the tool object with the alias key, never with the field name, or the field disappears without warning.&lt;/li>
&lt;li>Do not revalidate tool objects from a dump without asking for aliases.&lt;/li>
&lt;li>Count on marking a server as always loaded making session start-up wait for that server, and on a gateway listing live against all of its upstreams.&lt;/li>
&lt;li>Exclude in the client the native tools that compete, by bare name and not with a pattern in parentheses.&lt;/li>
&lt;li>Rewrite the descriptions of your own tools, and do not waste time renaming.&lt;/li>
&lt;li>Do not turn on proxy mode if you depend on the description overrides, because they are not applied there.&lt;/li>
&lt;li>Version the exposed tool set and its order as security configuration, with a hash of the description and the schema.&lt;/li>
&lt;li>If you use the gateway&amp;rsquo;s REST route for anything, know that it discards the &lt;code>_meta&lt;/code>.&lt;/li>
&lt;/ul>
&lt;h2 id="traps-and-things-that-are-not-what-they-look-like">Traps and things that are not what they look like&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>MCP&amp;rsquo;s &lt;code>priority&lt;/code> exists, and it is not about tools.&lt;/strong> It lives in the annotations for content, resources and prompts.&lt;/li>
&lt;li>&lt;strong>Tool annotations are hints and the schema itself says they are not reliable&lt;/strong> from untrusted servers.&lt;/li>
&lt;li>&lt;strong>The specification gives no precedence rule between servers&lt;/strong>, and it says besides that the server name is no use for disambiguating.&lt;/li>
&lt;li>&lt;strong>The tool set may indeed vary by authorisation&lt;/strong>, and that is the only blessed segmentation.&lt;/li>
&lt;li>&lt;strong>&lt;code>server/discover&lt;/code> does not return tools.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>A deny rule with parentheses does not remove the tool from the context&lt;/strong>, it only rejects the calls.&lt;/li>
&lt;li>&lt;strong>&lt;code>--tools&lt;/code> is a command-line flag&lt;/strong>, with no equivalent in the settings file.&lt;/li>
&lt;li>&lt;strong>Marking a server as always loaded delays start-up&lt;/strong> by up to five seconds per server.&lt;/li>
&lt;li>&lt;strong>&lt;code>Tool(meta=...)&lt;/code> does not populate the field&lt;/strong> and serialises it under the wrong key.&lt;/li>
&lt;li>&lt;strong>Proxy mode, search by permissions and the REST listing discard the &lt;code>_meta&lt;/code>.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Description overrides are skipped in proxy mode&lt;/strong> and are not applied on REST.&lt;/li>
&lt;li>&lt;strong>Proxy mode&amp;rsquo;s cap of five results can be raised by the operator&lt;/strong>, even though the client cannot ask for more.&lt;/li>
&lt;li>&lt;strong>The core tools of the proxy&amp;rsquo;s ranking are the only declarative priority in the stack&lt;/strong>, and they live in the mode that loses the &lt;code>_meta&lt;/code>.&lt;/li>
&lt;li>&lt;strong>You cannot force &amp;ldquo;any of mine from the gateway&amp;rdquo;&lt;/strong> in any API.&lt;/li>
&lt;li>&lt;strong>In OpenAI, the allowed list inside the MCP block and the identically named tool-choice type are different things.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Gemini does accept a per-tool allowlist&lt;/strong> on its remote MCP servers.&lt;/li>
&lt;li>&lt;strong>Reordering the catalogue repairs at most 23 % of the failures.&lt;/strong>&lt;/li>
&lt;/ul>
&lt;h2 id="closing">Closing&lt;/h2>
&lt;p>The starting question had a trap in it, and the trap is the word. When someone asks how you prioritise an MCP server, they are assuming a dial exists somewhere. It does not, and I have spent the whole article showing the places where it is not.&lt;/p>
&lt;p>What there is instead is poorer and more manageable: you remove what competes, you leave your own in front, and you write the description better. Three things, none of them elegant, all three effective in that order.&lt;/p>
&lt;p>What does seem to me worth taking away is the shape of the problem. The tool catalogue a model sees is security configuration in the full sense, because it determines what the agent can do and which system it is going to talk to. And today it is managed like a list of connections: things get added, it is not versioned, it is not signed, and nobody detects a change. The ecosystem has already solved that problem once, for skills, with hashes and tied approval. For tools, not yet.&lt;/p>
&lt;p>In the meantime, the asymmetry has to be built by hand, server by server, and knowing that it rests on the &lt;code>_meta&lt;/code> of an object that a library can empty without saying anything.&lt;/p>
&lt;h2 id="see-also">See also&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/litellm-mcp-gateway-second-front-door/">LiteLLM&amp;rsquo;s MCP gateway&lt;/a> — the article this one is a direct follow-up to, with the catalogue as cost and the ladder that gets corrected here.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/humans-agents-same-gateway-priority-traffic-separation/">Humans and agents on the same gateway&lt;/a> — the separation by key and team that credential-based filtering rests on.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/sizing-agents-decode-bottleneck/">Sizing for agents&lt;/a> — what it costs in proxy CPU to serialise the catalogue on every request.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/prefix-routing-what-litellm-does-not-do/">Prefix routing&lt;/a> — why touching the start of the prompt is paid for so dearly.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/completing-keycloak-for-mcp-protected-resource/">Completing Keycloak for MCP&lt;/a> — the protected resource side and audience validation.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/mcp-grows-up-authentication-keycloak/">When MCP grows: authentication with Keycloak&lt;/a> — the identity of your own MCP servers.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/mcp-deep-observability-opentelemetry/">MCP from the inside and its deep observability&lt;/a> — the protocol and its primitives.&lt;/li>
&lt;li>&lt;a href="https://blog.lo0.es/en/posts/technical-controls-ens-iso-42001-eu-ai-act-cross-mapping/">Technical controls for ENS, 42001 and the AI Act&lt;/a> — the framework where change management over the catalogue fits.&lt;/li>
&lt;/ul>
&lt;h2 id="sources">Sources&lt;/h2>
&lt;ul>
&lt;li>Model Context Protocol, herramientas en la revisión &lt;code>2026-07-28&lt;/code>: &lt;a href="https://modelcontextprotocol.io/specification/2026-07-28/server/tools">https://modelcontextprotocol.io/specification/2026-07-28/server/tools&lt;/a>.&lt;/li>
&lt;li>Model Context Protocol, &lt;code>server/discover&lt;/code>: &lt;a href="https://modelcontextprotocol.io/specification/2026-07-28/server/discover">https://modelcontextprotocol.io/specification/2026-07-28/server/discover&lt;/a>.&lt;/li>
&lt;li>Model Context Protocol, paginación: &lt;a href="https://modelcontextprotocol.io/specification/2026-07-28/server/utilities/pagination">https://modelcontextprotocol.io/specification/2026-07-28/server/utilities/pagination&lt;/a>.&lt;/li>
&lt;li>Model Context Protocol, caché de listados: &lt;a href="https://modelcontextprotocol.io/specification/2026-07-28/server/utilities/caching">https://modelcontextprotocol.io/specification/2026-07-28/server/utilities/caching&lt;/a>.&lt;/li>
&lt;li>Model Context Protocol, esquema fuente de la revisión: &lt;a href="https://github.com/modelcontextprotocol/modelcontextprotocol/blob/main/schema/2026-07-28/schema.ts">https://github.com/modelcontextprotocol/modelcontextprotocol/blob/main/schema/2026-07-28/schema.ts&lt;/a>.&lt;/li>
&lt;li>Model Context Protocol, extensión de habilidades: &lt;a href="https://modelcontextprotocol.io/extensions/skills/overview">https://modelcontextprotocol.io/extensions/skills/overview&lt;/a>.&lt;/li>
&lt;li>Model Context Protocol, grupo de interés sobre agrupación de primitivas: &lt;a href="https://modelcontextprotocol.io/community/interest-groups/primitive-grouping">https://modelcontextprotocol.io/community/interest-groups/primitive-grouping&lt;/a>.&lt;/li>
&lt;li>Anthropic, Claude Code y MCP (&lt;code>alwaysLoad&lt;/code>, búsqueda de herramientas): &lt;a href="https://code.claude.com/docs/en/mcp">https://code.claude.com/docs/en/mcp&lt;/a>.&lt;/li>
&lt;li>Anthropic, Claude Code, referencia de línea de comandos (&lt;code>--tools&lt;/code>): &lt;a href="https://code.claude.com/docs/en/cli-reference">https://code.claude.com/docs/en/cli-reference&lt;/a>.&lt;/li>
&lt;li>Anthropic, Claude Code, permisos (denegación por nombre pelado): &lt;a href="https://code.claude.com/docs/en/permissions">https://code.claude.com/docs/en/permissions&lt;/a>.&lt;/li>
&lt;li>Anthropic, conector MCP y conjuntos de herramientas: &lt;a href="https://platform.claude.com/docs/en/agents-and-tools/mcp-connector">https://platform.claude.com/docs/en/agents-and-tools/mcp-connector&lt;/a>.&lt;/li>
&lt;li>Anthropic, cambios de herramientas a mitad de conversación: &lt;a href="https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages">https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages&lt;/a>.&lt;/li>
&lt;li>Anthropic, búsqueda de herramientas: &lt;a href="https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool">https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool&lt;/a>.&lt;/li>
&lt;li>OpenAI, MCP y conectores (lista de herramientas permitidas): &lt;a href="https://developers.openai.com/api/docs/guides/tools-connectors-mcp">https://developers.openai.com/api/docs/guides/tools-connectors-mcp&lt;/a>.&lt;/li>
&lt;li>OpenAI, búsqueda de herramientas y caché: &lt;a href="https://developers.openai.com/api/docs/guides/tools-tool-search">https://developers.openai.com/api/docs/guides/tools-tool-search&lt;/a>.&lt;/li>
&lt;li>Google, llamada a funciones en la Interactions API: &lt;a href="https://ai.google.dev/gemini-api/docs/function-calling">https://ai.google.dev/gemini-api/docs/function-calling&lt;/a>.&lt;/li>
&lt;li>vLLM, llamada a herramientas y decodificación restringida: &lt;a href="https://docs.vllm.ai/en/latest/features/tool_calling.html">https://docs.vllm.ai/en/latest/features/tool_calling.html&lt;/a>.&lt;/li>
&lt;li>vLLM, propuesta de decodificación guiada por regiones y gramáticas de herramienta: &lt;a href="https://github.com/vllm-project/vllm/issues/39848">https://github.com/vllm-project/vllm/issues/39848&lt;/a>.&lt;/li>
&lt;li>LiteLLM, código del gateway MCP: &lt;code>litellm/proxy/_experimental/mcp_server/server.py&lt;/code>, &lt;code>mcp_server_manager.py&lt;/code>, &lt;code>rest_endpoints.py&lt;/code>, &lt;code>tool_search.py&lt;/code>, &lt;code>litellm/experimental_mcp_client/tools.py&lt;/code>: &lt;a href="https://github.com/BerriAI/litellm">https://github.com/BerriAI/litellm&lt;/a>.&lt;/li>
&lt;li>SDK de Python de MCP, modelo &lt;code>Tool&lt;/code> y serialización por alias: &lt;a href="https://github.com/modelcontextprotocol/python-sdk">https://github.com/modelcontextprotocol/python-sdk&lt;/a>.&lt;/li>
&lt;li>&lt;em>BiasBusters&lt;/em>, sesgo de selección entre herramientas equivalentes y peso de la descripción frente al nombre: &lt;a href="https://arxiv.org/html/2510.00307v2">https://arxiv.org/html/2510.00307v2&lt;/a>.&lt;/li>
&lt;li>&lt;em>Looking Is Not Picking&lt;/em>, atención frente a lectura final y techo de las intervenciones sobre el prompt: &lt;a href="https://arxiv.org/html/2606.16364v1">https://arxiv.org/html/2606.16364v1&lt;/a>.&lt;/li>
&lt;li>&lt;em>A Single Rewrite Suffices&lt;/em>, reescritura de descripciones y generalización al conjunto recuperado: &lt;a href="https://arxiv.org/html/2606.30775">https://arxiv.org/html/2606.30775&lt;/a>.&lt;/li>
&lt;li>&lt;em>ToolFlood&lt;/em>, saturación adversaria de la capa de recuperación: &lt;a href="https://arxiv.org/html/2603.13950">https://arxiv.org/html/2603.13950&lt;/a>.&lt;/li>
&lt;/ul></description></item></channel></rss>