# Bouwen van een "Advocaat GPT" voor uw blog - Deel 6: Lokale LLM-integratie

<!--category-- AI, LLM, LLamaSharp, GGUF, C#, AI-Article, mostlylucid.blogllm -->
<datetime class="hidden">2025-11-12T22:45</datetime>

WAARSCHUWING: Deze zijn ontwerp-Posts die 'onthuld'.

Het is waarschijnlijk veel van wat hieronder zal niet werken; IK genereren deze als how-to voor MIJ en dan doen alle stappen en krijg de sample app werken...Je bent stiekem en ze gezien! ze zullen waarschijnlijk klaar zijn medio december.

<img src="https://media1.tenor.com/m/_rQc7PIEqwQAAAAd/cat-hello-cat-peek.gif" height="300px" />
## Inleiding

Welkom in deel 6![We hebben de complete infrastructuur gebouwd - inslikken pijpleiding (](/blog/building-a-lawyer-gpt-for-your-blog-part4)Deel 4[), Windows client (](/blog/building-a-lawyer-gpt-for-your-blog-part5)Deel 5[), inbeddingen en vectorzoeken (](/blog/building-a-lawyer-gpt-for-your-blog-part3)Deel 3[), en GPU setup (](/blog/building-a-lawyer-gpt-for-your-blog-part2)Deel 2



Nu komt het spannende deel: het integreren van een lokale LLM om daadwerkelijk schrijven suggesties te genereren.

[TOC]

## OPMERKING: Dit maakt deel uit van mijn experimenten met AI (ondersteund opstellen) + mijn eigen bewerking.

Dezelfde stem, hetzelfde pragmatisme, gewoon snellere vingers.

### Hier maken we eindelijk het "AI" deel van "AI schrijfassistent" werk.

We zullen grote taalmodellen lokaal uitvoeren op uw A4000 GPU, het genereren van context-aware suggesties op basis van uw vorige blog berichten.
|--------|-----------|-------------------|
| **Waarom lokale LLM?** | ✅ Complete | ❌ Data sent to third party |
| **Voordat we naar binnen duiken, laten we begrijpen waarom we modellen lokaal draaien in plaats van OpenAI's API te gebruiken.** | ✅ Free after setup | ❌ Per-token pricing |
| **Lokale vergelijking vs. API** | ✅ <1 second | ⚠️ Network dependent |
| **Lokaal LLM API (OpenAI, enz.)** | ✅ Full control | ❌ Limited |
| **Privacy** | ✅ Any GGUF model | ❌ Provider's models only |
| **Kosten** | ✅ Works offline | ❌ Requires internet |
| **Matigheid** | ❌ Complex | ✅ Simple |

Aanpassen

## Modelkeuze

Offline

```mermaid
graph TB
    A[C# Application] --> B{Integration Method}

    B --> C[LLamaSharp]
    B --> D[ONNX Runtime]
    B --> E[TorchSharp]
    B --> F[HTTP API]

    C --> G[llama.cpp bindings]
    G --> H[GGUF Models]

    D --> I[ONNX Models]
    I --> J[Limited Model Support]

    E --> K[PyTorch Models]
    K --> L[Complex Setup]

    F --> M[External Process]
    M --> N[Ollama, LM Studio]

    class C recommended
    class G,H llamaSharp

    classDef recommended stroke:#333,stroke-width:4px
    classDef llamaSharp stroke:#333,stroke-width:2px
```

**Instellen**

Voor een schrijfassistent, privacy en kosten.

- We willen niet dat blogontwerpen worden verzonden naar externe API's, en per-token prijzen kloppen snel voor een dagelijks schrijven tool.
- LLM-integratieopties voor C#
- Er zijn verschillende manieren om LLM's uit te voeren in C#:
- Mijn keuze: LLamaSharp
- Waarom?

## Native C# bindingen voor lama.cpp (fastest inference library)

### Ondersteunt GGUF-formaat (moderne, gequantiseerde modellen)

[CUDA-versnelling ingebouwd](https://github.com/ggerganov/ggml/blob/master/docs/gguf.md)Actieve ontwikkeling en grote gemeenschap

```mermaid
graph LR
    A[Original Model<br/>Llama 2 7B<br/>~28GB float32] --> B[Quantization]

    B --> C[Q4_K_M<br/>~4.1GB<br/>4-bit]
    B --> D[Q5_K_M<br/>~4.8GB<br/>5-bit]
    B --> E[Q6_K<br/>~5.5GB<br/>6-bit]
    B --> F[Q8_0<br/>~7.2GB<br/>8-bit]

    C --> G[Fast, Lower Quality]
    D --> H[Balanced]
    E --> I[Higher Quality]
    F --> J[Near Original]

    class A original
    class C,D quantized
    class H recommended

    classDef original stroke:#333,stroke-width:2px
    classDef quantized stroke:#333,stroke-width:2px
    classDef recommended stroke:#333,stroke-width:2px
```

**Werkt met Llama, Mistral, Phi, Gemma, en nog veel meer**:

- Modelformaten en kwantisering begrijpen
- GGUF-formaat
- GGUF
- (GPT-Generated Unified Format) is de standaard voor het efficiënt uitvoeren van LLM's.

### Kwantisering uitgelegd

Origineel model: 32-bits praalwagens (zeer groot, zeer nauwkeurig)
|-------|---------------|------------|-----------|------------|------------|---------|
| **Q4: 4-bit gehele getallen (75% kleiner, minimaal kwaliteitsverlies)** | 2.3GB | ~4GB | ✅ Easy | ✅ Easy | ✅ Easy | ⭐⭐⭐ Good |
| **Q5/Q6: Zoete plek voor de meeste gebruiks gevallen** | 4.1GB | ~6GB | ✅ Tight | ✅ Good | ✅ Easy | ⭐⭐⭐ Good |
| **Q8: Bijna-originele kwaliteit, nog 4x kleiner** | 4.1GB | ~6GB | ✅ Tight | ✅ Good | ✅ Easy | ⭐⭐⭐⭐ Better |
| **Modelselectie door Hardware** | 4.1GB | ~6GB | ✅ Tight | ✅ Good | ✅ Easy | ⭐⭐⭐⭐ Better |
| **Maat (Q4_K_M) VRAM-gebruik Past 8GB? Past 12GB? Past 16GB? Kwaliteit** | 4.7GB | ~7GB | ⚠️ Very Tight | ✅ Good | ✅ Easy | ⭐⭐⭐⭐⭐ Best |
| **Phi-3 Mini (3.8B)** | 7.4GB | ~10GB | ❌ No | ⚠️ Tight | ✅ Good | ⭐⭐⭐⭐ Better |

**Llama 2 7B**

- **Mistral 7B**Gemma 7B**Llama 3 8B**Llama 2 13B**Aanbevelingen van de GPU:**8GB VRAM
- **: Beginnen met**: **Mistral 7B**of**Phi-3 Mini**(Safest)
- **12GB VRAM**: **Llama 3 8B**(beste kwaliteit) of**Mistral 7B**
- **(sneller)**16GB VRAM (mijn installatie)

**Llama 3 8B**: **[of probeer](https://mistral.ai/)**13B modellen**[Alleen CPU](https://ai.meta.com/llama/)**: Elk model werkt, net veel langzamer (begin met Phi-3 Mini voor snelheid)

- Mijn aanbeveling
- Mistral 7B
- (laatste versie) of
- Llama 3

## 8B

Uitstekende kwaliteit voor technisch schrijven

### Werkt in alle GPU maten

1. Snel genoeg voor interactief gebruik[Goed in het volgen van instructies](https://huggingface.co/models)
2. Downloaden van modellen`"mistral 7b gguf"`
3. Modellen worden verspreid op Hugging Face.

**We gebruiken quantized GGUF versies.**GGUF-modellen vinden

- [Ga naar](https://huggingface.co/TheBloke/Mistral-7B-Instruct-v0.2-GGUF)
- [Gezicht knuffelen](https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGUF)
- [Zoeken:](https://huggingface.co/QuantFactory/Meta-Llama-3-8B-Instruct-GGUF)

### Zoek naar de kwantificaties van TheBloke (meest populair)

```bash
# Install huggingface-cli
pip install huggingface-hub

# Download Mistral 7B Q5_K_M (recommended)
huggingface-cli download TheBloke/Mistral-7B-Instruct-v0.2-GGUF \
    mistral-7b-instruct-v0.2.Q5_K_M.gguf \
    --local-dir C:\models\mistral-7b \
    --local-dir-use-symlinks False
```

Directe links

1. (Kwantificeringen van TheBloke):
2. Mistral-7B-Instruct-v0.2-GGUF`mistral-7b-instruct-v0.2.Q5_K_M.gguf`Llama-2-7B-Chat-GGUF
3. Llama-3-8B-Instruct-GGUF
4. Specifieke kwantisatie downloaden`C:\models\mistral-7b\`

## Of handmatig downloaden:

### Klik op "Bestanden en versies" tabblad

```bash
cd Mostlylucid.BlogLLM.Core
dotnet add package LLamaSharp  # Latest version
dotnet add package LLamaSharp.Backend.Cuda12  # Latest, matching CUDA version
```

**Zoeken**

- `[LLamaSharp](https://github.com/SciSharp/LLamaSharp)`(~4,8 GB)
- `LLamaSharp.Backend.Cuda12` - [Klik op downloaden](https://developer.nvidia.com/cuda-toolkit)Opslaan naar

### LLamaSharp instellen

Pakket NuGet installeren

```csharp
using LLama;
using LLama.Common;

// Check if CUDA is available
bool cudaAvailable = NativeLibraryConfig.Instance.CudaEnabled;
Console.WriteLine($"CUDA Available: {cudaAvailable}");
```

Waarom twee pakjes?`false`- Kernbibliotheek

1. CUDA
2. `LLamaSharp.Backend.Cuda12`12 binaries voor GPU acceleratie
3. CUDA-backend verifiëren

## LLamaSharp zal CUDA automatisch detecteren als het correct is geïnstalleerd.

### Als

```csharp
using LLama;
using LLama.Common;

namespace Mostlylucid.BlogLLM.Core.Services
{
    public class ModelParameters
    {
        public string ModelPath { get; set; } = string.Empty;
        public int ContextSize { get; set; } = 4096;  // Context window
        public int GpuLayerCount { get; set; } = 35;  // Layers on GPU (35 = all for 7B)
        public int Seed { get; set; } = 1337;  // For reproducibility
        public float Temperature { get; set; } = 0.7f;  // Creativity (0.0 = deterministic, 1.0 = creative)
        public float TopP { get; set; } = 0.9f;  // Nucleus sampling
        public int MaxTokens { get; set; } = 500;  // Max generation length
    }
}
```

**, controleer:**:

- **CUDA 12.x geïnstalleerd (deel 2)**pakket geïnstalleerd
  
  - PATH bevat CUDA bin directory
  - Bouwen van de LLM-dienst

- **Modelparameters**Uitleg van parameters
  
  - Contextgrootte
  - : Hoeveel tekst het model tegelijk kan "zien"
  - 4096 tokens ≈ 3000 woorden

- **Groter = meer context maar langzamer en meer VRAM**GpuLayerCount
  
  - : Hoeveel transformatorlagen draaien op GPU
  - 7B modellen hebben ~32 lagen
  - 35 = zet alles op GPU (snelste)

- **Lagere waarden = minder VRAM maar langzamer gebruiken**Temperatuur
  
  - : Controleert randomness
  - 0.0 = altijd kiezen meest waarschijnlijke token (saai, repetitief)

### 0,7 = goed saldo (onze standaard)

```csharp
using LLama;
using LLama.Common;
using Microsoft.Extensions.Logging;

namespace Mostlylucid.BlogLLM.Core.Services
{
    public interface ILlmService
    {
        Task<string> GenerateAsync(string prompt, CancellationToken cancellationToken = default);
        Task<string> GenerateWithContextAsync(string prompt, List<string> contextChunks, CancellationToken cancellationToken = default);
    }

    public class LlmService : ILlmService, IDisposable
    {
        private readonly LLamaWeights _model;
        private readonly LLamaContext _context;
        private readonly ILogger<LlmService> _logger;
        private readonly ModelParameters _parameters;

        public LlmService(ModelParameters parameters, ILogger<LlmService> logger)
        {
            _parameters = parameters;
            _logger = logger;

            _logger.LogInformation("Loading model from {ModelPath}", parameters.ModelPath);

            // Configure model parameters
            var modelParams = new ModelParams(parameters.ModelPath)
            {
                ContextSize = (uint)parameters.ContextSize,
                GpuLayerCount = parameters.GpuLayerCount,
                Seed = (uint)parameters.Seed,
                UseMemoryLock = true,  // Keep model in RAM
                UseMemorymap = true    // Memory-map the model file
            };

            // Load model
            _model = LLamaWeights.LoadFromFile(modelParams);
            _context = _model.CreateContext(modelParams);

            _logger.LogInformation("Model loaded successfully. VRAM used: ~{VRAM}GB",
                EstimateVRAMUsage(parameters.GpuLayerCount));
        }

        public async Task<string> GenerateAsync(string prompt, CancellationToken cancellationToken = default)
        {
            var executor = new InteractiveExecutor(_context);

            var inferenceParams = new InferenceParams
            {
                Temperature = _parameters.Temperature,
                TopP = _parameters.TopP,
                MaxTokens = _parameters.MaxTokens,
                AntiPrompts = new[] { "\n\nUser:", "###" }  // Stop generation at these
            };

            var result = new StringBuilder();

            _logger.LogInformation("Generating response for prompt: {Prompt}", TruncateForLog(prompt));

            await foreach (var token in executor.InferAsync(prompt, inferenceParams, cancellationToken))
            {
                result.Append(token);
            }

            var response = result.ToString().Trim();
            _logger.LogInformation("Generated {Tokens} tokens", CountTokens(response));

            return response;
        }

        public async Task<string> GenerateWithContextAsync(
            string prompt,
            List<string> contextChunks,
            CancellationToken cancellationToken = default)
        {
            // Build prompt with retrieved context
            var fullPrompt = BuildContextualPrompt(prompt, contextChunks);

            _logger.LogInformation("Context chunks: {Count}, Total prompt tokens: ~{Tokens}",
                contextChunks.Count, CountTokens(fullPrompt));

            return await GenerateAsync(fullPrompt, cancellationToken);
        }

        private string BuildContextualPrompt(string userPrompt, List<string> contextChunks)
        {
            var sb = new StringBuilder();

            sb.AppendLine("You are a helpful writing assistant for a technical blog.");
            sb.AppendLine("Use the following excerpts from past blog posts as context:");
            sb.AppendLine();

            for (int i = 0; i < contextChunks.Count; i++)
            {
                sb.AppendLine($"--- Context {i + 1} ---");
                sb.AppendLine(contextChunks[i]);
                sb.AppendLine();
            }

            sb.AppendLine("---");
            sb.AppendLine();
            sb.AppendLine("Based on the context above, help with the following:");
            sb.AppendLine(userPrompt);
            sb.AppendLine();
            sb.AppendLine("Response:");

            return sb.ToString();
        }

        private int CountTokens(string text)
        {
            // Rough estimate: 1 token ≈ 4 characters
            return text.Length / 4;
        }

        private string TruncateForLog(string text, int maxLength = 100)
        {
            if (text.Length <= maxLength) return text;
            return text.Substring(0, maxLength) + "...";
        }

        private double EstimateVRAMUsage(int gpuLayers)
        {
            // Rough estimate for 7B model
            return (gpuLayers / 35.0) * 6.0;  // ~6GB for full 7B model
        }

        public void Dispose()
        {
            _context?.Dispose();
            _model?.Dispose();
        }
    }
}
```

**1.0+ = zeer creatief (kan nonsensisch zijn)**:

1. **TopP**: Nucleus-bemonstering
2. **0,9 = rekening houden met penningen die 90% van de waarschijnlijkheidsmassa uitmaken**Voorkomt bemonstering van zeer onwaarschijnlijke tokens
3. **Implementatie van LLM-diensten**Hoe het werkt
4. **Model Laden**: laadt GGUF model in VRAM met behulp van gespecificeerde parameters
5. **Interactieve uitvoerder**: LLamaSharp's uitvoering modus voor chat-achtige interacties

### InferAsync

```csharp
using Microsoft.Extensions.Logging;

class Program
{
    static async Task Main(string[] args)
    {
        // Setup logging
        var loggerFactory = LoggerFactory.Create(builder => builder.AddConsole());
        var logger = loggerFactory.CreateLogger<LlmService>();

        // Configure model
        var parameters = new ModelParameters
        {
            ModelPath = @"C:\models\mistral-7b\mistral-7b-instruct-v0.2.Q5_K_M.gguf",
            ContextSize = 4096,
            GpuLayerCount = 35,
            Temperature = 0.7f,
            MaxTokens = 200
        };

        // Create service
        using var llmService = new LlmService(parameters, logger);

        // Test simple generation
        Console.WriteLine("=== Test 1: Simple Generation ===\n");
        var response1 = await llmService.GenerateAsync(
            "Explain what Docker Compose is in 2-3 sentences."
        );
        Console.WriteLine(response1);
        Console.WriteLine("\n");

        // Test with context
        Console.WriteLine("=== Test 2: Generation with Context ===\n");
        var context = new List<string>
        {
            "Docker Compose is a tool for defining and running multi-container Docker applications. With Compose, you use a YAML file to configure your application's services.",
            "In development, Docker Compose makes it easy to spin up all dependencies (databases, caches, etc.) with one command: docker-compose up."
        };

        var response2 = await llmService.GenerateWithContextAsync(
            "Write an introduction paragraph for a blog post about using Docker Compose for development dependencies.",
            context
        );
        Console.WriteLine(response2);
    }
}
```

**: Streams tokens als ze worden gegenereerd (real-time output)**:

```
=== Test 1: Simple Generation ===

Docker Compose is a tool that allows you to define and run multi-container Docker applications using a simple YAML configuration file. It simplifies the process of managing multiple containers, networking, and volumes, making it ideal for development environments.

=== Test 2: Generation with Context ===

If you've ever found yourself juggling multiple terminal windows to start databases, caches, and other services for local development, Docker Compose is about to become your new best friend. This powerful tool lets you define your entire development environment in a single YAML file and spin everything up with one command. In this post, we'll explore how to leverage Docker Compose to manage all your development dependencies, making your local setup reproducible, shareable, and incredibly easy to manage.
```

Contextgebouw

## : Combineert de gebruiker prompt met opgehaalde blog brokken

Antiprompts

### : Stopt generatie op bepaalde strings (voorkomt gerammel)

```csharp
namespace Mostlylucid.BlogLLM.Client.Services
{
    public class SuggestionService : ISuggestionService
    {
        private readonly BatchEmbeddingService _embeddingService;
        private readonly QdrantVectorStore _vectorStore;
        private readonly ILlmService _llmService;  // NEW

        public SuggestionService(
            BatchEmbeddingService embeddingService,
            QdrantVectorStore vectorStore,
            ILlmService llmService)  // NEW
        {
            _embeddingService = embeddingService;
            _vectorStore = vectorStore;
            _llmService = llmService;
        }

        public async Task<string> GenerateAiSuggestionAsync(
            string currentText,
            List<SimilarPost> context)
        {
            // Extract text from similar posts
            var contextChunks = context
                .Take(3)  // Top 3 most similar
                .Select(p => p.FullText)
                .ToList();

            // Determine what type of suggestion to generate
            var prompt = DeterminePromptType(currentText);

            // Generate suggestion
            var suggestion = await _llmService.GenerateWithContextAsync(
                prompt,
                contextChunks
            );

            return suggestion;
        }

        private string DeterminePromptType(string currentText)
        {
            // Analyze what user is writing
            var lines = currentText.Split('\n');
            var lastLine = lines.LastOrDefault(l => !string.IsNullOrWhiteSpace(l)) ?? "";

            // Is user starting a new section?
            if (lastLine.StartsWith("## "))
            {
                return "Suggest 3-5 bullet points for what this section could cover.";
            }

            // Is user writing code?
            if (lastLine.Contains("```"))
            {
                return "Suggest relevant code examples that might be useful here.";
            }

            // Is user writing an introduction?
            if (currentText.Length < 500 && currentText.Contains("## Introduction"))
            {
                return "Suggest 2-3 sentences to continue this introduction based on similar posts.";
            }

            // Default: continue current thought
            return "Suggest 1-2 sentences to continue the current paragraph in a natural way.";
        }
    }
}
```

### Testen van de Dienst

```csharp
public partial class SuggestionsViewModel : ViewModelBase
{
    [RelayCommand]
    private async Task RegenerateSuggestion()
    {
        IsGenerating = true;
        AiSuggestion = "Generating...";

        try
        {
            var currentText = GetCurrentEditorText();  // From messaging
            var suggestion = await _suggestionService.GenerateAiSuggestionAsync(
                currentText,
                SimilarPosts.ToList()
            );

            AiSuggestion = suggestion;
        }
        catch (Exception ex)
        {
            AiSuggestion = $"Error: {ex.Message}";
        }
        finally
        {
            IsGenerating = false;
        }
    }
}
```

## Verwachte output

### Geweldig!

Het model werkt en genereert coherente, contextbewuste tekst.

```csharp
public class LlmServiceFactory
{
    private static LlmService? _instance;
    private static readonly object _lock = new();

    public static LlmService GetInstance(ModelParameters parameters, ILogger<LlmService> logger)
    {
        if (_instance == null)
        {
            lock (_lock)
            {
                if (_instance == null)
                {
                    _instance = new LlmService(parameters, logger);
                }
            }
        }

        return _instance;
    }
}
```

### Integratie met Suggesties Panel

Laten we nu LLM-generatie integreren in onze Windows-client van deel 5.

```csharp
public class StatefulLlmService
{
    private readonly InferenceParams _defaultParams;
    private string _cachedPromptPrefix = string.Empty;

    public async Task<string> GenerateWithPrefixAsync(string prefix, string newPrompt)
    {
        // If prefix matches cached, reuse KV cache
        if (prefix == _cachedPromptPrefix)
        {
            // Only process new tokens
            return await GenerateAsync(newPrompt);
        }

        // Process entire prompt and cache
        _cachedPromptPrefix = prefix;
        return await GenerateAsync(prefix + newPrompt);
    }
}
```

SuggestieService bijwerken

### Update SuggestiesBekijkModel

Optimalisatie van de prestaties

```csharp
public async Task<List<string>> GenerateBatchAsync(List<string> prompts)
{
    var results = new List<string>();

    foreach (var prompt in prompts)
    {
        // With KV cache reuse, subsequent prompts are faster
        results.Add(await GenerateAsync(prompt));
    }

    return results;
}
```

## Model Caching

Houd het model geladen tussen de verzoeken:

### KV Cache Hergebruik

```csharp
private string PromptContinueWriting(string currentText, List<string> context)
{
    return $@"You are a technical blog writing assistant.

Here are excerpts from similar blog posts:
{string.Join("\n\n", context.Select((c, i) => $"--- Post {i + 1} ---\n{c}"))}

Current draft:
{currentText}

Task: Suggest 2-3 sentences to naturally continue the current paragraph.
Keep the same technical depth and casual, pragmatic tone.

Suggestion:";
}
```

### LLamaSharp ondersteunt KV cache hergebruik voor snellere volgende generaties:

```csharp
private string PromptSectionStructure(string sectionTitle, List<string> context)
{
    return $@"You are a technical blog writing assistant.

Similar sections from past posts:
{string.Join("\n\n", context)}

New section: {sectionTitle}

Task: Suggest 4-6 bullet points for what this section should cover.
Format as a markdown list.

Bullets:";
}
```

### Dit is vooral handig voor onze use case - de context brokken blijven hetzelfde, alleen de vraag van de gebruiker verandert.

```csharp
private string PromptCodeExample(string description, List<string> context)
{
    return $@"You are a C# coding assistant.

Relevant code from past posts:
{string.Join("\n\n", context)}

Task: {description}

Provide a clean, well-commented C# code example.

Code:";
}
```

## Loting

Voor meerdere suggesties, batch ze:

### Prompt Engineering voor Schrijfhulp

```csharp
public class VramMonitor
{
    [DllImport("nvml.dll")]
    private static extern int nvmlDeviceGetMemoryInfo(IntPtr device, ref NvmlMemory memory);

    [StructLayout(LayoutKind.Sequential)]
    public struct NvmlMemory
    {
        public ulong Total;
        public ulong Free;
        public ulong Used;
    }

    public static (ulong used, ulong total) GetVramUsage()
    {
        // Simplified - actual implementation needs proper NVML initialization
        var memory = new NvmlMemory();
        // nvmlDeviceGetMemoryInfo(device, ref memory);

        return (memory.Used / 1024 / 1024, memory.Total / 1024 / 1024);  // Convert to MB
    }
}
```

### Goede prompts = goede output.

```csharp
public class LlmServiceWithUnload : IDisposable
{
    private LlmService? _service;
    private readonly Timer _unloadTimer;
    private DateTime _lastUsed;

    public LlmServiceWithUnload()
    {
        _unloadTimer = new Timer(CheckForUnload, null, TimeSpan.FromMinutes(1), TimeSpan.FromMinutes(1));
    }

    private void CheckForUnload(object? state)
    {
        if (_service != null && (DateTime.Now - _lastUsed) > TimeSpan.FromMinutes(10))
        {
            _service.Dispose();
            _service = null;
            GC.Collect();
            Console.WriteLine("Model unloaded due to inactivity");
        }
    }

    public async Task<string> GenerateAsync(string prompt)
    {
        _lastUsed = DateTime.Now;

        if (_service == null)
        {
            // Reload model
            _service = CreateService();
        }

        return await _service.GenerateAsync(prompt);
    }
}
```

## Hier zijn sjablonen voor verschillende scenario's:

Schrijven verder

```csharp
public async Task<string> GenerateWithRetryAsync(string prompt, int maxRetries = 3)
{
    for (int i = 0; i < maxRetries; i++)
    {
        try
        {
            return await GenerateAsync(prompt);
        }
        catch (OutOfMemoryException)
        {
            _logger.LogWarning("OOM error, reducing max tokens");
            _parameters.MaxTokens = Math.Max(100, _parameters.MaxTokens / 2);
        }
        catch (Exception ex)
        {
            _logger.LogError(ex, "Generation failed, attempt {Attempt}/{Max}", i + 1, maxRetries);

            if (i == maxRetries - 1) throw;

            await Task.Delay(1000 * (i + 1));  // Exponential backoff
        }
    }

    throw new Exception("Generation failed after retries");
}
```

## Deelstructuur voorstellen

Codevoorbeeld

1. ✅ Chose [Geheugenbeheer](https://github.com/SciSharp/LLamaSharp)Met grote modellen is geheugenbeheer cruciaal.
2. ✅ Understood [VRAM-gebruik monitoren](https://github.com/ggerganov/ggml/blob/master/docs/gguf.md)Model uitladen als het niet actief is
3. ✅ Selected appropriate model ([Fout bij omgaan](https://mistral.ai/) / [LLM's kunnen op onverwachte manieren falen.](https://ai.meta.com/llama/)Handvat sierlijk:
4. ✅ Implemented LlmService with CUDA acceleration
5. ✅ Integrated with Windows client for suggestions
6. ✅ Implemented prompt engineering for writing tasks
7. ✅ Added performance optimizations (caching, batching)
8. ✅ Handled memory management and errors

## Samenvatting

We hebben de lokale LLM-inferentie succesvol geïntegreerd:**[LLamaSharp](/blog/building-a-lawyer-gpt-for-your-blog-part7)**voor C# integratie

- GGUF-formaat
- en quantisering
- Mistral 7B
- Llama 3
- 8B)
- Wat is het volgende?
- In

Deel 7: Content Generation & Prompt Engineering

## , we zullen ons richten op de volledige content generation pipeline:

- [Geavanceerde prompt engineering technieken](/blog/building-a-lawyer-gpt-for-your-blog-part1)
- [Multi-turn conversatie voor iteratieve verfijning](/blog/building-a-lawyer-gpt-for-your-blog-part2)
- [Strategieën voor het beheer van het contextvenster](/blog/building-a-lawyer-gpt-for-your-blog-part3)
- [Kwaliteitsbeoordeling en filtering](/blog/building-a-lawyer-gpt-for-your-blog-part4)
- [Stijlverbintenis handhaving](/blog/building-a-lawyer-gpt-for-your-blog-part5)
- **Aanmaak van codeblok voor verwerking**Gebruikspatronen in de reële wereld
- [We maken het systeem eigenlijk nuttig voor het dagelijks schrijven van blogs!](/blog/building-a-lawyer-gpt-for-your-blog-part7)
- [Serienavigatie](/blog/building-a-lawyer-gpt-for-your-blog-part8)

## Deel 1: Inleiding en architectuur

- [Deel 2: GPU-instellingen & CUDA in C#](https://scisharp.github.io/LLamaSharp/)
- [Deel 3: Inbeddingen en vectordatabases begrijpen](https://github.com/ggerganov/ggml/blob/master/docs/gguf.md)
- [Deel 4: Bouwen van de Ingestie Pijpleiding](https://huggingface.co/TheBloke)
- [Deel 5: De Windows Client](https://github.com/ggerganov/llama.cpp)

Deel 6: Lokale LLM-integratie[(dit bericht)](/blog/building-a-lawyer-gpt-for-your-blog-part7)!