Costruire un "GPT avvocato" per il tuo blog - Parte 6: Integrazione LLM locale (Italiano (Italian))

Costruire un "GPT avvocato" per il tuo blog - Parte 6: Integrazione LLM locale

Wednesday, 12 November 2025

//

18 minute read

ATTENZIONE: Questi sono i progetti di POST che sono "scaparnati."

E 'probabile che gran parte di ciò che è qui sotto non funzionerà; Io generare questi come come come-per me e poi fare tutti i passi e ottenere il campione di lavoro app ... Sei stato subdolo e li hai visti! saranno probabilmente pronti a metà dicembre.

## Introduzione

Benvenuti alla Parte 6Abbiamo costruito l'infrastruttura completa - pipeline di ingestione (Parte 4), Windows client (Parte 5), incorporazioni e ricerca vettoriale (Parte 3), e la configurazione della GPU (Parte 2

Ora arriva la parte emozionante: integrare un LLM locale per generare effettivamente suggerimenti di scrittura.

NOTA: Questo fa parte dei miei esperimenti con l'AI (elaborazione assistita) + il mio editing.

Stessa voce, stesso pragmatismo, solo dita piu' veloci.

E' qui che finalmente facciamo parte del lavoro "AI" di "AI writing assistant."

Eseguiremo grandi modelli di lingua localmente sulla tua GPU A4000, generando suggerimenti contestuali basati sui post del tuo blog passato. |--------|-----------|-------------------| | Perché LLM locale? | ✅ Complete | ❌ Data sent to third party | | Prima di immergerci, cerchiamo di capire perché stiamo eseguendo i modelli localmente invece di utilizzare l'API di OpenAI. | ✅ Free after setup | ❌ Per-token pricing | | Confronto locale vs API | ✅ <1 second | ⚠️ Network dependent | | | Aspetto | LLM locale | API (OpenAI, ecc.) | | ✅ Full control | ❌ Limited | | Privacy | ✅ Any GGUF model | ❌ Provider's models only | | Costo | ✅ Works offline | ❌ Requires internet | | Latenza | ❌ Complex | ✅ Simple |

Personalizzazione

Scelta del modello

Fuori rete

graph TB
    A[C# Application] --> B{Integration Method}

    B --> C[LLamaSharp]
    B --> D[ONNX Runtime]
    B --> E[TorchSharp]
    B --> F[HTTP API]

    C --> G[llama.cpp bindings]
    G --> H[GGUF Models]

    D --> I[ONNX Models]
    I --> J[Limited Model Support]

    E --> K[PyTorch Models]
    K --> L[Complex Setup]

    F --> M[External Process]
    M --> N[Ollama, LM Studio]

    class C recommended
    class G,H llamaSharp

    classDef recommended stroke:#333,stroke-width:4px
    classDef llamaSharp stroke:#333,stroke-width:2px

Configurazione

Per un assistente di scrittura, la privacy e i costi sono importanti.

  • Non vogliamo che le bozze di blog inviate ad API esterne, e prezzi per-token aggiunge veloce per uno strumento di scrittura quotidiana.
  • Opzioni di integrazione LLM per C#
  • Ci sono diversi modi per eseguire LLM in C#:
  • La mia scelta: LLamaSharp
  • Perché?

Legami C# nativi per lama.cpp (libreria di inferenza più veloce)

Supporta il formato GGUF (moderni, quantizzati)

Accelerazione CUDA integrataSviluppo attivo e grande comunità

graph LR
    A[Original Model<br/>Llama 2 7B<br/>~28GB float32] --> B[Quantization]

    B --> C[Q4_K_M<br/>~4.1GB<br/>4-bit]
    B --> D[Q5_K_M<br/>~4.8GB<br/>5-bit]
    B --> E[Q6_K<br/>~5.5GB<br/>6-bit]
    B --> F[Q8_0<br/>~7.2GB<br/>8-bit]

    C --> G[Fast, Lower Quality]
    D --> H[Balanced]
    E --> I[Higher Quality]
    F --> J[Near Original]

    class A original
    class C,D quantized
    class H recommended

    classDef original stroke:#333,stroke-width:2px
    classDef quantized stroke:#333,stroke-width:2px
    classDef recommended stroke:#333,stroke-width:2px

Funziona con Llama, Mistral, Phi, Gemma, e altro ancora:

  • Comprendere i formati del modello e la quantizzazione
  • Formato GGUF
  • GGUFCity name (optional, probably does not need a translation)
  • (GPT-Generated Unified Format) è lo standard per l'esecuzione efficiente di LLM.

Quantizzazione spiegata

Modello originale: float a 32 bit (molto grande, molto preciso) |-------|---------------|------------|-----------|------------|------------|---------| | Q4: interi a 4 bit (75% più piccolo, perdita minima di qualità) | 2.3GB | ~4GB | ✅ Easy | ✅ Easy | ✅ Easy | ⭐⭐⭐ Good | | Q5/Q6: Sweet spot per la maggior parte dei casi di utilizzo | 4.1GB | ~6GB | ✅ Tight | ✅ Good | ✅ Easy | ⭐⭐⭐ Good | | Q8: Qualità quasi originale, ancora 4x più piccola | 4.1GB | ~6GB | ✅ Tight | ✅ Good | ✅ Easy | ⭐⭐⭐⭐ Better | | Selezione del modello per hardware | 4.1GB | ~6GB | ✅ Tight | ✅ Good | ✅ Easy | ⭐⭐⭐⭐ Better | | | Model | Size (Q4_K_M) | VRAM Usage | Fits 8GB? | Fits 12GB? | Fits 16GB? | Quality | | 4.7GB | ~7GB | ⚠️ Very Tight | ✅ Good | ✅ Easy | ⭐⭐⭐⭐⭐ Best | | Phi-3 Mini (3.8B) | 7.4GB | ~10GB | ❌ No | ⚠️ Tight | ✅ Good | ⭐⭐⭐⭐ Better |

Llama 2 7B

  • Mistral 7BGemma 7BLlama 3 8BLlama 2 13B**Raccomandazioni della GPU:**VRAM da 8GB
  • : Inizia con: Mistral 7BoppurePhi-3 MiniCity name (optional, probably does not need a translation)(Safesto)
  • VRAM da 12GB: Llama 3 8B(migliore qualità) oMistral 7B
  • **(Più veloce)**VRAM da 16GB (la mia configurazione)

Llama 3 8B: o provaModelli 13BSolo CPU: Qualsiasi modello funziona, solo molto più lento (iniziare con Phi-3 Mini per la velocità)

  • La mia raccomandazione
  • Mistral 7B
  • (ultima versione) oppure
  • Llama 3City name (optional, probably does not need a translation)

8B

Ottima qualità per la scrittura tecnica

Funziona su tutte le dimensioni della GPU

  1. Abbastanza veloce per l'uso interattivoBuono a seguire le istruzioni
  2. Scaricamento modelli"mistral 7b gguf"
  3. I modelli sono distribuiti su Hugging Face.

**Useremo le versioni quantizzate del GGUF.**Trovare modelli GGUF

Cerca le quantizzazioni di TheBloke (più popolari)

# Install huggingface-cli
pip install huggingface-hub

# Download Mistral 7B Q5_K_M (recommended)
huggingface-cli download TheBloke/Mistral-7B-Instruct-v0.2-GGUF \
    mistral-7b-instruct-v0.2.Q5_K_M.gguf \
    --local-dir C:\models\mistral-7b \
    --local-dir-use-symlinks False

Collegamenti diretti

  1. (La quantizzazione del Bloke):
  2. Mistral-7B-Instruct-v0.2-GUFmistral-7b-instruct-v0.2.Q5_K_M.ggufLlama-2-7B-Chat-GGUF
  3. Llama-3-8B-Instruct-GGUF
  4. Scarica Quantizzazione specificaC:\models\mistral-7b\

O scaricare manualmente:

Clicca sulla scheda "File e versioni"

cd Mostlylucid.BlogLLM.Core
dotnet add package LLamaSharp  # Latest version
dotnet add package LLamaSharp.Backend.Cuda12  # Latest, matching CUDA version

Trova

  • [LLamaSharp](https://github.com/SciSharp/LLamaSharp)(~4.8GB)
  • LLamaSharp.Backend.Cuda12 - Fare clic sul downloadSalva a

Configurazione LLamaSharp

Installa il pacchetto NuGet

using LLama;
using LLama.Common;

// Check if CUDA is available
bool cudaAvailable = NativeLibraryConfig.Instance.CudaEnabled;
Console.WriteLine($"CUDA Available: {cudaAvailable}");

Perche' due pacchetti?false- Libreria centrale

  1. CUDACity name (optional, probably does not need a translation)
  2. LLamaSharp.Backend.Cuda1212 binari per l'accelerazione della GPU
  3. Verificare il backend CUDA

LLamaSharp rileverà automaticamente CUDA se installato correttamente.

Se

using LLama;
using LLama.Common;

namespace Mostlylucid.BlogLLM.Core.Services
{
    public class ModelParameters
    {
        public string ModelPath { get; set; } = string.Empty;
        public int ContextSize { get; set; } = 4096;  // Context window
        public int GpuLayerCount { get; set; } = 35;  // Layers on GPU (35 = all for 7B)
        public int Seed { get; set; } = 1337;  // For reproducibility
        public float Temperature { get; set; } = 0.7f;  // Creativity (0.0 = deterministic, 1.0 = creative)
        public float TopP { get; set; } = 0.9f;  // Nucleus sampling
        public int MaxTokens { get; set; } = 500;  // Max generation length
    }
}

, controllo::

  • **CUDA 12.x installato (parte 2)**pacchetto installato

    • PATH include la directory CUDA bin
    • Costruzione del servizio LLM
  • Parametri del modelloSpiegazioni dei parametri

    • ContextSize
    • : Quanto testo il modello può "vedere" in una sola volta
    • 4096 gettoni ~ 3000 parole
  • Più grande = più contesto ma più lento e più VRAMGpuLayerCountCity name (optional, probably does not need a translation)

    • : Quanti livelli di trasformatore vengono eseguiti sulla GPU
    • I modelli 7B hanno ~32 livelli
    • 35 = mettere tutto sulla GPU (più veloce)
  • Valori inferiori = usare meno VRAM ma più lentamenteTemperatura

    • : Controlla la casualità
    • 0.0 = sceglie sempre il token più probabile (borioso, ripetitivo)

0,7 = buon saldo (il nostro valore predefinito)

using LLama;
using LLama.Common;
using Microsoft.Extensions.Logging;

namespace Mostlylucid.BlogLLM.Core.Services
{
    public interface ILlmService
    {
        Task<string> GenerateAsync(string prompt, CancellationToken cancellationToken = default);
        Task<string> GenerateWithContextAsync(string prompt, List<string> contextChunks, CancellationToken cancellationToken = default);
    }

    public class LlmService : ILlmService, IDisposable
    {
        private readonly LLamaWeights _model;
        private readonly LLamaContext _context;
        private readonly ILogger<LlmService> _logger;
        private readonly ModelParameters _parameters;

        public LlmService(ModelParameters parameters, ILogger<LlmService> logger)
        {
            _parameters = parameters;
            _logger = logger;

            _logger.LogInformation("Loading model from {ModelPath}", parameters.ModelPath);

            // Configure model parameters
            var modelParams = new ModelParams(parameters.ModelPath)
            {
                ContextSize = (uint)parameters.ContextSize,
                GpuLayerCount = parameters.GpuLayerCount,
                Seed = (uint)parameters.Seed,
                UseMemoryLock = true,  // Keep model in RAM
                UseMemorymap = true    // Memory-map the model file
            };

            // Load model
            _model = LLamaWeights.LoadFromFile(modelParams);
            _context = _model.CreateContext(modelParams);

            _logger.LogInformation("Model loaded successfully. VRAM used: ~{VRAM}GB",
                EstimateVRAMUsage(parameters.GpuLayerCount));
        }

        public async Task<string> GenerateAsync(string prompt, CancellationToken cancellationToken = default)
        {
            var executor = new InteractiveExecutor(_context);

            var inferenceParams = new InferenceParams
            {
                Temperature = _parameters.Temperature,
                TopP = _parameters.TopP,
                MaxTokens = _parameters.MaxTokens,
                AntiPrompts = new[] { "\n\nUser:", "###" }  // Stop generation at these
            };

            var result = new StringBuilder();

            _logger.LogInformation("Generating response for prompt: {Prompt}", TruncateForLog(prompt));

            await foreach (var token in executor.InferAsync(prompt, inferenceParams, cancellationToken))
            {
                result.Append(token);
            }

            var response = result.ToString().Trim();
            _logger.LogInformation("Generated {Tokens} tokens", CountTokens(response));

            return response;
        }

        public async Task<string> GenerateWithContextAsync(
            string prompt,
            List<string> contextChunks,
            CancellationToken cancellationToken = default)
        {
            // Build prompt with retrieved context
            var fullPrompt = BuildContextualPrompt(prompt, contextChunks);

            _logger.LogInformation("Context chunks: {Count}, Total prompt tokens: ~{Tokens}",
                contextChunks.Count, CountTokens(fullPrompt));

            return await GenerateAsync(fullPrompt, cancellationToken);
        }

        private string BuildContextualPrompt(string userPrompt, List<string> contextChunks)
        {
            var sb = new StringBuilder();

            sb.AppendLine("You are a helpful writing assistant for a technical blog.");
            sb.AppendLine("Use the following excerpts from past blog posts as context:");
            sb.AppendLine();

            for (int i = 0; i < contextChunks.Count; i++)
            {
                sb.AppendLine($"--- Context {i + 1} ---");
                sb.AppendLine(contextChunks[i]);
                sb.AppendLine();
            }

            sb.AppendLine("---");
            sb.AppendLine();
            sb.AppendLine("Based on the context above, help with the following:");
            sb.AppendLine(userPrompt);
            sb.AppendLine();
            sb.AppendLine("Response:");

            return sb.ToString();
        }

        private int CountTokens(string text)
        {
            // Rough estimate: 1 token ≈ 4 characters
            return text.Length / 4;
        }

        private string TruncateForLog(string text, int maxLength = 100)
        {
            if (text.Length <= maxLength) return text;
            return text.Substring(0, maxLength) + "...";
        }

        private double EstimateVRAMUsage(int gpuLayers)
        {
            // Rough estimate for 7B model
            return (gpuLayers / 35.0) * 6.0;  // ~6GB for full 7B model
        }

        public void Dispose()
        {
            _context?.Dispose();
            _model?.Dispose();
        }
    }
}

1,0+ = molto creativo (può essere insensato):

  1. TopP: campionamento del nucleo
  2. 0,9 = considerare i gettoni che costituiscono il 90% della massa di probabilitàImpedisce il campionamento da gettoni molto improbabili
  3. Implementazione dei servizi LLMCome funziona
  4. Caricamento modello: Carica il modello GGUF in VRAM utilizzando parametri specificati
  5. InteractiveExecutor: Modalità di esecuzione di LLamaSharp per interazioni simili a chat

InferAsyncCity name (optional, probably does not need a translation)

using Microsoft.Extensions.Logging;

class Program
{
    static async Task Main(string[] args)
    {
        // Setup logging
        var loggerFactory = LoggerFactory.Create(builder => builder.AddConsole());
        var logger = loggerFactory.CreateLogger<LlmService>();

        // Configure model
        var parameters = new ModelParameters
        {
            ModelPath = @"C:\models\mistral-7b\mistral-7b-instruct-v0.2.Q5_K_M.gguf",
            ContextSize = 4096,
            GpuLayerCount = 35,
            Temperature = 0.7f,
            MaxTokens = 200
        };

        // Create service
        using var llmService = new LlmService(parameters, logger);

        // Test simple generation
        Console.WriteLine("=== Test 1: Simple Generation ===\n");
        var response1 = await llmService.GenerateAsync(
            "Explain what Docker Compose is in 2-3 sentences."
        );
        Console.WriteLine(response1);
        Console.WriteLine("\n");

        // Test with context
        Console.WriteLine("=== Test 2: Generation with Context ===\n");
        var context = new List<string>
        {
            "Docker Compose is a tool for defining and running multi-container Docker applications. With Compose, you use a YAML file to configure your application's services.",
            "In development, Docker Compose makes it easy to spin up all dependencies (databases, caches, etc.) with one command: docker-compose up."
        };

        var response2 = await llmService.GenerateWithContextAsync(
            "Write an introduction paragraph for a blog post about using Docker Compose for development dependencies.",
            context
        );
        Console.WriteLine(response2);
    }
}

: Streams token generati (output in tempo reale):

=== Test 1: Simple Generation ===

Docker Compose is a tool that allows you to define and run multi-container Docker applications using a simple YAML configuration file. It simplifies the process of managing multiple containers, networking, and volumes, making it ideal for development environments.

=== Test 2: Generation with Context ===

If you've ever found yourself juggling multiple terminal windows to start databases, caches, and other services for local development, Docker Compose is about to become your new best friend. This powerful tool lets you define your entire development environment in a single YAML file and spin everything up with one command. In this post, we'll explore how to leverage Docker Compose to manage all your development dependencies, making your local setup reproducible, shareable, and incredibly easy to manage.

Costruzione di contesto

: Combina il prompt utente con pezzi di blog recuperati

Antiprompt

: Ferma la generazione a certe corde (previene il disordine)

namespace Mostlylucid.BlogLLM.Client.Services
{
    public class SuggestionService : ISuggestionService
    {
        private readonly BatchEmbeddingService _embeddingService;
        private readonly QdrantVectorStore _vectorStore;
        private readonly ILlmService _llmService;  // NEW

        public SuggestionService(
            BatchEmbeddingService embeddingService,
            QdrantVectorStore vectorStore,
            ILlmService llmService)  // NEW
        {
            _embeddingService = embeddingService;
            _vectorStore = vectorStore;
            _llmService = llmService;
        }

        public async Task<string> GenerateAiSuggestionAsync(
            string currentText,
            List<SimilarPost> context)
        {
            // Extract text from similar posts
            var contextChunks = context
                .Take(3)  // Top 3 most similar
                .Select(p => p.FullText)
                .ToList();

            // Determine what type of suggestion to generate
            var prompt = DeterminePromptType(currentText);

            // Generate suggestion
            var suggestion = await _llmService.GenerateWithContextAsync(
                prompt,
                contextChunks
            );

            return suggestion;
        }

        private string DeterminePromptType(string currentText)
        {
            // Analyze what user is writing
            var lines = currentText.Split('\n');
            var lastLine = lines.LastOrDefault(l => !string.IsNullOrWhiteSpace(l)) ?? "";

            // Is user starting a new section?
            if (lastLine.StartsWith("## "))
            {
                return "Suggest 3-5 bullet points for what this section could cover.";
            }

            // Is user writing code?
            if (lastLine.Contains("```"))
            {
                return "Suggest relevant code examples that might be useful here.";
            }

            // Is user writing an introduction?
            if (currentText.Length < 500 && currentText.Contains("## Introduction"))
            {
                return "Suggest 2-3 sentences to continue this introduction based on similar posts.";
            }

            // Default: continue current thought
            return "Suggest 1-2 sentences to continue the current paragraph in a natural way.";
        }
    }
}

Verifica del servizio

public partial class SuggestionsViewModel : ViewModelBase
{
    [RelayCommand]
    private async Task RegenerateSuggestion()
    {
        IsGenerating = true;
        AiSuggestion = "Generating...";

        try
        {
            var currentText = GetCurrentEditorText();  // From messaging
            var suggestion = await _suggestionService.GenerateAiSuggestionAsync(
                currentText,
                SimilarPosts.ToList()
            );

            AiSuggestion = suggestion;
        }
        catch (Exception ex)
        {
            AiSuggestion = $"Error: {ex.Message}";
        }
        finally
        {
            IsGenerating = false;
        }
    }
}

Uscita prevista

Fantastico!

Il modello sta lavorando e generando un testo coerente e contestuale.

public class LlmServiceFactory
{
    private static LlmService? _instance;
    private static readonly object _lock = new();

    public static LlmService GetInstance(ModelParameters parameters, ILogger<LlmService> logger)
    {
        if (_instance == null)
        {
            lock (_lock)
            {
                if (_instance == null)
                {
                    _instance = new LlmService(parameters, logger);
                }
            }
        }

        return _instance;
    }
}

Integrare con il Pannello Suggerimenti

Ora integriamo la generazione LLM nel nostro client Windows dalla parte 5.

public class StatefulLlmService
{
    private readonly InferenceParams _defaultParams;
    private string _cachedPromptPrefix = string.Empty;

    public async Task<string> GenerateWithPrefixAsync(string prefix, string newPrompt)
    {
        // If prefix matches cached, reuse KV cache
        if (prefix == _cachedPromptPrefix)
        {
            // Only process new tokens
            return await GenerateAsync(newPrompt);
        }

        // Process entire prompt and cache
        _cachedPromptPrefix = prefix;
        return await GenerateAsync(prefix + newPrompt);
    }
}

Aggiorna SuggerimentoServizio

Aggiorna suggerimentiViewModel

Ottimizzazione delle prestazioni

public async Task<List<string>> GenerateBatchAsync(List<string> prompts)
{
    var results = new List<string>();

    foreach (var prompt in prompts)
    {
        // With KV cache reuse, subsequent prompts are faster
        results.Add(await GenerateAsync(prompt));
    }

    return results;
}

Modello Caching

Mantieni il modello caricato tra le richieste:

Riutilizzo della cache di KV

private string PromptContinueWriting(string currentText, List<string> context)
{
    return $@"You are a technical blog writing assistant.

Here are excerpts from similar blog posts:
{string.Join("\n\n", context.Select((c, i) => $"--- Post {i + 1} ---\n{c}"))}

Current draft:
{currentText}

Task: Suggest 2-3 sentences to naturally continue the current paragraph.
Keep the same technical depth and casual, pragmatic tone.

Suggestion:";
}

LLamaSharp supporta il riutilizzo della cache KV per generazioni successive più veloci:

private string PromptSectionStructure(string sectionTitle, List<string> context)
{
    return $@"You are a technical blog writing assistant.

Similar sections from past posts:
{string.Join("\n\n", context)}

New section: {sectionTitle}

Task: Suggest 4-6 bullet points for what this section should cover.
Format as a markdown list.

Bullets:";
}

Questo è particolarmente utile per il nostro caso d'uso - i pezzi di contesto rimangono gli stessi, solo la domanda dell'utente cambia.

private string PromptCodeExample(string description, List<string> context)
{
    return $@"You are a C# coding assistant.

Relevant code from past posts:
{string.Join("\n\n", context)}

Task: {description}

Provide a clean, well-commented C# code example.

Code:";
}

Lotto

Per i suggerimenti multipli, batch loro:

Prompt Engineering for Writing Assistance

public class VramMonitor
{
    [DllImport("nvml.dll")]
    private static extern int nvmlDeviceGetMemoryInfo(IntPtr device, ref NvmlMemory memory);

    [StructLayout(LayoutKind.Sequential)]
    public struct NvmlMemory
    {
        public ulong Total;
        public ulong Free;
        public ulong Used;
    }

    public static (ulong used, ulong total) GetVramUsage()
    {
        // Simplified - actual implementation needs proper NVML initialization
        var memory = new NvmlMemory();
        // nvmlDeviceGetMemoryInfo(device, ref memory);

        return (memory.Used / 1024 / 1024, memory.Total / 1024 / 1024);  // Convert to MB
    }
}

Buoni suggerimenti = buona uscita.

public class LlmServiceWithUnload : IDisposable
{
    private LlmService? _service;
    private readonly Timer _unloadTimer;
    private DateTime _lastUsed;

    public LlmServiceWithUnload()
    {
        _unloadTimer = new Timer(CheckForUnload, null, TimeSpan.FromMinutes(1), TimeSpan.FromMinutes(1));
    }

    private void CheckForUnload(object? state)
    {
        if (_service != null && (DateTime.Now - _lastUsed) > TimeSpan.FromMinutes(10))
        {
            _service.Dispose();
            _service = null;
            GC.Collect();
            Console.WriteLine("Model unloaded due to inactivity");
        }
    }

    public async Task<string> GenerateAsync(string prompt)
    {
        _lastUsed = DateTime.Now;

        if (_service == null)
        {
            // Reload model
            _service = CreateService();
        }

        return await _service.GenerateAsync(prompt);
    }
}

Ecco i modelli per diversi scenari:

Continua a scrivere

public async Task<string> GenerateWithRetryAsync(string prompt, int maxRetries = 3)
{
    for (int i = 0; i < maxRetries; i++)
    {
        try
        {
            return await GenerateAsync(prompt);
        }
        catch (OutOfMemoryException)
        {
            _logger.LogWarning("OOM error, reducing max tokens");
            _parameters.MaxTokens = Math.Max(100, _parameters.MaxTokens / 2);
        }
        catch (Exception ex)
        {
            _logger.LogError(ex, "Generation failed, attempt {Attempt}/{Max}", i + 1, maxRetries);

            if (i == maxRetries - 1) throw;

            await Task.Delay(1000 * (i + 1));  // Exponential backoff
        }
    }

    throw new Exception("Generation failed after retries");
}

Suggerisci la struttura della sezione

Esempio di codice

  1. ✅ Chose Gestione della memoriaCon i grandi modelli, la gestione della memoria è cruciale.
  2. ✅ Understood Monitorare l'utilizzo della VRAMScarica il modello quando inattivo
  3. ✅ Selected appropriate model (Gestione degli errori / I LLM possono fallire in modi inaspettati.Gestisci con grazia:
  4. ✅ Implemented LlmService with CUDA acceleration
  5. ✅ Integrated with Windows client for suggestions
  6. ✅ Implemented prompt engineering for writing tasks
  7. ✅ Added performance optimizations (caching, batching)
  8. ✅ Handled memory management and errors

Sommario

Abbiamo integrato con successo l'inferenza LLM locale:**LLamaSharp**per l'integrazione C#

  • Formato GGUF
  • e quantizzazione
  • Mistral 7B
  • Llama 3City name (optional, probably does not need a translation)
  • 8B)
  • Cosa c'e' dopo?
  • Dentro

Parte 7: Generazione dei contenuti e Prompt Engineering

, ci concentreremo sulla completa pipeline di generazione di contenuti:

Parte 1: Introduzione e architettura

Parte 6: Integrazione LLM locale(questo posto)!

Finding related posts...
logo

© 2026 Scott Galloway — Unlicense — All content and source code on this site is free to use, copy, modify, and sell.