{"id":6768,"date":"2026-09-05T13:07:39","date_gmt":"2026-09-05T13:07:39","guid":{"rendered":"https:\/\/justfineinfotech.com\/shopify-introduces-gisting-compressing-llm-system-prompts-into-learned-tokens\/"},"modified":"2026-09-05T13:07:39","modified_gmt":"2026-09-05T13:07:39","slug":"shopify-introduces-gisting-compressing-llm-system-prompts-into-learned-tokens","status":"publish","type":"post","link":"https:\/\/justfineinfotech.com\/fr\/shopify-introduces-gisting-compressing-llm-system-prompts-into-learned-tokens\/","title":{"rendered":"Shopify Introduces Gisting: Compressing LLM System Prompts into Learned Tokens"},"content":{"rendered":"<p>InfoQ HomepageNewsShopify Introduces Gisting: Compressing LLM System Prompts into Learned Tokens<\/p>\n<p>AI, ML &amp; Data Engineering<br \/>\nBeyond the PR: The New Control Plane for Agentic Software Delivery (Webinar Sept 17th)<\/p>\n<p>Shopify&#8217;s engineering introduced Gisting, a novel technique for compressing long LLM prompts into a smaller set of learned &#8220;gist&#8221; tokens, improving throughput and reducing inference cost<\/p>\n<p>Spotify emphasizes that replacing lengthy text for concise gist tokens at inference time reduces end-to-end latency, drops infrastructure costs, and boosts token throughput without modifying the model&#8217;s core weights<\/p>\n<p>The company says that gisting reduced the Sidekick GraphQL agent\u2019s system prompt from about 6000 tokens to 1500 gist tokens without sacrificing prediction quality. This implies a 4:1 reduction in context size:<\/p>\n<blockquote>\n<p>At 350 requests per minute (RPM), the median time to first token (TTFT) dropped from 438ms to 354ms, the median end-to-end request latency dropped from 6.8s to 4.2s, and throughput rose from 20.2 to 23.4 queries per second (QPS)<\/p>\n<\/blockquote>\n<p>These improved metrics allowed Spotify to reduce the number of allocated GPUs<\/p>\n<p>Gisting is based on a technique pioneered in a 2022 paper, <a href=\"https:\/\/arxiv.org\/abs\/2210.03162\" rel=\"nofollow noopener\" target=\"_blank\">&#8220;<em>Prompt Compression and Contrastive Conditioning for Controllability and Toxicity Reduction in Language Models<\/em>&#8220;<\/a>, and consists of a two-step process to learn the embeddings of the new compressed gist token. In a first pass, the <em>teacher pass<\/em>, the model is run with the real prompt to derive the <em>teacher logits<\/em> of the response. In the <em>student pass<\/em>, the model is run with the gist tokens to derive the student logits. Finally, the gist are trained to minimize the KL divergence between the teacher logits and the student logits, that is until the student&#8217;s predictions closely match the teacher&#8217;s.<\/p>\n<blockquote>\n<p>When training finishes, we write the gist embeddings straight into the model&#8217;s embedding matrix, and register the new gist tokens as special tokens in the model\u2019s tokenizer. The model loads and runs like any other at inference time: no custom attention mask, extra encoder, or special serving <a href=\"https:\/\/justfineinfotech.com\/fr\/how-the-creator-economy-is-rewriting-the-path-to-stardom\/\" title=\"Comment l&#039;\u00e9conomie des cr\u00e9ateurs red\u00e9finit le chemin vers la c\u00e9l\u00e9brit\u00e9\">path<\/a><\/p>\n<\/blockquote>\n<p>The key advantage of gisting is that the model does not process a conventional summary of the original prompt, but rather a learned representation designed to make the LLM to behave as close as possible to how it would if it had seen the original prompt<\/p>\n<p>Gisting can reduce latency and increase throughput. In Shopify&#8217;s case, Time to First Token (TTFT) dropped from 438ms to 354ms, and end-to-end latency fell from 6.8s to 4.2s. At the same time, queries per second (QPS) increased from 20.2 to 23.4, allowing engineering teams to scale down overall GPU allocation<\/p>\n<p>As a final note, Shopify also emphasizes that gisting is complementary to other optimization techniques, such as <a href=\"https:\/\/handbook.modular.com\/inference-optimization\/prefix-caching\/\" rel=\"nofollow noopener\" target=\"_blank\">prefix caching<\/a>. Prefix caching avoids recomputing the KV tensors for cached prompt sequences, but the model must still process those cached tensors during the decoding phase. Gisting further reduces this overhead by replacing a long prompt with a shorter sequence of learned gist tokens. The two optimizations therefore compound, and Shopify uses them together.<\/p>\n<p>There is much more to gisting than can be covered here. Make sure to read the original article if you are interested in the full details, which covers topics such as the role of autosearch in tuning the Gisting process and other implementation details that significantly affect performance<\/p>\n<div style=\"clear:both;margin:30px 0 15px 0\">\n<p>\n    <strong>En rapport:<\/strong><br \/>\n    &lt;a href=&quot;https:\/\/yoursite.com\/automation-training-benin\/&quot; title=&quot;<a href=\"https:\/\/justfineinfotech.com\/fr\/all-resources-to-be-utilised-to-make-kp-digital-economy-hub-minister\/\" title=\"All resources to be utilised to make KP digital economy hub: minister\">Num\u00e9rique<\/a> Formation en automatisation au B\u00e9nin\u00a0: 5 comp\u00e9tences gagnantes que les employeurs recherchent en 2026<br \/>\n      Formation en automatisation num\u00e9rique au B\u00e9nin\u00a0: 5 comp\u00e9tences cl\u00e9s recherch\u00e9es par les employeurs en 2026<br \/>\n    <\/a>\n  <\/p>\n<p>\n    &lt;a href=&quot;https:\/\/yoursite.com\/automation-africa\/&quot; title=&quot;<a href=\"https:\/\/justfineinfotech.com\/fr\/data-shows-83-of-festive-whatsapp-orders-came-from-first-time-buyers\/\" title=\"Data shows 83% of festive Whatsapp orders came from first-time buyers\">WhatsApp<\/a> Marketing Automation Africa : 6 erreurs dangereuses commises par les marques au Nig\u00e9ria\u201d&gt;<br \/>\n      Automatisation du marketing WhatsApp en Afrique\u00a0: 6 erreurs dangereuses commises par les marques au Nig\u00e9ria<br \/>\n    <\/a>\n  <\/p>\n<\/div>\n<div style=\"clear:both;margin:30px 0;padding:25px;background:#f8f9fc;border:1px solid #ddd;border-radius:8px;text-align:center\">\n<h3>Vous souhaitez apprendre cela de mani\u00e8re pratique ?<\/h3>\n<p>Rejoindre <strong>Justfine Infotech<\/strong> et d\u00e9velopper de v\u00e9ritables comp\u00e9tences num\u00e9riques en IA, automatisation, d\u00e9veloppement web, marketing digital, bureautique, e-commerce, travail ind\u00e9pendant et cybers\u00e9curit\u00e9.<\/p>\n<p><strong>Programmes disponibles :<\/strong><br \/>\n  Certificat de 6 semaines \u2022 Certificat professionnel de 3 mois \u2022 Dipl\u00f4me de 6 mois \u2022 Dipl\u00f4me professionnel complet<\/p>\n<p><strong>WhatsApp :<\/strong><br \/>\n  +229 01 57 57 99 15<br \/>\n  +229 01 66 68 11 60<\/p>\n<p><a href=\"https:\/\/api.whatsapp.com\/send?phone=2348132690270&amp;text=Hello\" target=\"_blank\" rel=\"noopener\">Inscrivez-vous d\u00e8s maintenant<\/a><\/p>\n<\/div>\n<p class=\"ani-source\">Source: <a href=\"https:\/\/www.infoq.com\/news\/2026\/09\/spotify-gisting-llm-performance\/\" target=\"_blank\" rel=\"nofollow noopener\">www.infoq.com<\/a><\/p>","protected":false},"excerpt":{"rendered":"<p>InfoQ HomepageNewsShopify Introduces Gisting: Compressing LLM System Prompts into Learned Tokens<\/p>","protected":false},"author":1,"featured_media":6770,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":""},"categories":[70],"tags":[1597,1596,102,149,952],"class_list":["post-6768","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-e-commerce-digital-payments","tag-compressing","tag-gisting","tag-introduces","tag-shopify","tag-system"],"_links":{"self":[{"href":"https:\/\/justfineinfotech.com\/fr\/wp-json\/wp\/v2\/posts\/6768","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/justfineinfotech.com\/fr\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/justfineinfotech.com\/fr\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/justfineinfotech.com\/fr\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/justfineinfotech.com\/fr\/wp-json\/wp\/v2\/comments?post=6768"}],"version-history":[{"count":1,"href":"https:\/\/justfineinfotech.com\/fr\/wp-json\/wp\/v2\/posts\/6768\/revisions"}],"predecessor-version":[{"id":6769,"href":"https:\/\/justfineinfotech.com\/fr\/wp-json\/wp\/v2\/posts\/6768\/revisions\/6769"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/justfineinfotech.com\/fr\/wp-json\/wp\/v2\/media\/6770"}],"wp:attachment":[{"href":"https:\/\/justfineinfotech.com\/fr\/wp-json\/wp\/v2\/media?parent=6768"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/justfineinfotech.com\/fr\/wp-json\/wp\/v2\/categories?post=6768"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/justfineinfotech.com\/fr\/wp-json\/wp\/v2\/tags?post=6768"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}