# Construyendo un "Abogado GPT" para su blog - Parte 6: Integración local de LLM

<!--category-- AI, LLM, LLamaSharp, GGUF, C#, AI-Article, mostlylucid.blogllm -->
<datetime class="hidden">2025-11-12T22:45</datetime>

ATENCIÓN: Estos son los proyectos de mensajes que se han "espaciado".

Es probable que gran parte de lo que está a continuación no funcione; Yo genera estos como cómo para MÍ y luego hacer todos los pasos y hacer que la aplicación de muestra funcione...¡Has sido astuto y los has visto!Probablemente estarán listos a mediados de diciembre.

<img src="https://media1.tenor.com/m/_rQc7PIEqwQAAAAd/cat-hello-cat-peek.gif" height="300px" />
## Introducción

¡Bienvenidos a la Parte 6![Hemos construido la infraestructura completa - tubería de ingestión (](/blog/building-a-lawyer-gpt-for-your-blog-part4)Parte 4[), cliente de Windows (](/blog/building-a-lawyer-gpt-for-your-blog-part5)Parte 5[), incrustaciones y búsqueda de vectores (](/blog/building-a-lawyer-gpt-for-your-blog-part3)Parte 3[), y configuración de la GPU (](/blog/building-a-lawyer-gpt-for-your-blog-part2)Parte 2



Ahora viene la parte emocionante: la integración de un LLM local para generar sugerencias de escritura.

[TOC]

## NOTA: Esto es parte de mis experimentos con IA (redacción asistida) + mi propia edición.

La misma voz, el mismo pragmatismo; sólo los dedos más rápidos.

### Aquí es donde finalmente hacemos que la "AI" sea parte de "AI writing assistant" funcione.

Ejecutaremos grandes modelos de idiomas localmente en tu GPU A4000, generando sugerencias de contexto basadas en tus posts de blog anteriores.
|--------|-----------|-------------------|
| **¿Por qué LLM local?** | ✅ Complete | ❌ Data sent to third party |
| **Antes de sumergirnos, vamos a entender por qué estamos ejecutando modelos localmente en lugar de usar la API de OpenAI.** | ✅ Free after setup | ❌ Per-token pricing |
| **Comparación entre API y local** | ✅ <1 second | ⚠️ Network dependent |
| **Aspect  LLM local  API (OpenAI, etc.)** | ✅ Full control | ❌ Limited |
| **Privacidad** | ✅ Any GGUF model | ❌ Provider's models only |
| **Costo** | ✅ Works offline | ❌ Requires internet |
| **Latencia** | ❌ Complex | ✅ Simple |

Personalización

## Elección del modelo

Desconectado

```mermaid
graph TB
    A[C# Application] --> B{Integration Method}

    B --> C[LLamaSharp]
    B --> D[ONNX Runtime]
    B --> E[TorchSharp]
    B --> F[HTTP API]

    C --> G[llama.cpp bindings]
    G --> H[GGUF Models]

    D --> I[ONNX Models]
    I --> J[Limited Model Support]

    E --> K[PyTorch Models]
    K --> L[Complex Setup]

    F --> M[External Process]
    M --> N[Ollama, LM Studio]

    class C recommended
    class G,H llamaSharp

    classDef recommended stroke:#333,stroke-width:4px
    classDef llamaSharp stroke:#333,stroke-width:2px
```

**Configuración**

Para un asistente de escritura, la privacidad y el costo de la materia.

- No queremos que los borradores de blog se envíen a APIs externas, y los precios por token se suman rápidamente para una herramienta de escritura diaria.
- Opciones de integración LLM para C#
- Hay varias maneras de ejecutar LLMs en C#:
- Mi elección: LLAMASharp
- ¿Por qué?

## Encuadernaciones nativas de C# para llama.cpp (biblioteca de inferencia más rápida)

### Soporta formato GGUF (moderno, modelos cuantificados)

[Aceleración CUDA incorporada](https://github.com/ggerganov/ggml/blob/master/docs/gguf.md)Desarrollo activo y gran comunidad

```mermaid
graph LR
    A[Original Model<br/>Llama 2 7B<br/>~28GB float32] --> B[Quantization]

    B --> C[Q4_K_M<br/>~4.1GB<br/>4-bit]
    B --> D[Q5_K_M<br/>~4.8GB<br/>5-bit]
    B --> E[Q6_K<br/>~5.5GB<br/>6-bit]
    B --> F[Q8_0<br/>~7.2GB<br/>8-bit]

    C --> G[Fast, Lower Quality]
    D --> H[Balanced]
    E --> I[Higher Quality]
    F --> J[Near Original]

    class A original
    class C,D quantized
    class H recommended

    classDef original stroke:#333,stroke-width:2px
    classDef quantized stroke:#333,stroke-width:2px
    classDef recommended stroke:#333,stroke-width:2px
```

**Trabaja con Llama, Mistral, Phi, Gemma y más**:

- Entender los formatos de modelos y la cuantificación
- Formato GGUF
- GGUF
- (GPT-Genered Unified Format) es el estándar para ejecutar LLMs de manera eficiente.

### Cuantificación explicada

Modelo original: flotadores de 32 bits (muy grandes, muy precisos)
|-------|---------------|------------|-----------|------------|------------|---------|
| **Q4: enteros de 4 bits (75% más pequeños, pérdida de calidad mínima)** | 2.3GB | ~4GB | ✅ Easy | ✅ Easy | ✅ Easy | ⭐⭐⭐ Good |
| **Q5/Q6: Lugar dulce para la mayoría de los casos de uso** | 4.1GB | ~6GB | ✅ Tight | ✅ Good | ✅ Easy | ⭐⭐⭐ Good |
| **Q8: Calidad casi original, aún 4 veces más pequeña** | 4.1GB | ~6GB | ✅ Tight | ✅ Good | ✅ Easy | ⭐⭐⭐⭐ Better |
| **Selección de modelos por hardware** | 4.1GB | ~6GB | ✅ Tight | ✅ Good | ✅ Easy | ⭐⭐⭐⭐ Better |
| **Modelo  Tamaño (Q4_K_M)  Uso VRAM  Se ajusta a 8GB?  Se ajusta a 12GB?  Se ajusta a 16GB?  Calidad** | 4.7GB | ~7GB | ⚠️ Very Tight | ✅ Good | ✅ Easy | ⭐⭐⭐⭐⭐ Best |
| **Phi-3 Mini (3,8B)** | 7.4GB | ~10GB | ❌ No | ⚠️ Tight | ✅ Good | ⭐⭐⭐⭐ Better |

**Llama 2 7B**

- **Mistral 7B**Gemma 7B**Llama 3 8B**Llama 2 13B**Recomendaciones de la GPU:**8GB VRAM
- **: Comience con**: **Mistral 7B**o**Phi-3 Mini**(más seguro)
- **12GB VRAM**: **Llama 3 8B**(mejor calidad) o**Mistral 7B**
- **(Más rápido)**16GB VRAM (mi configuración)

**Llama 3 8B**: **[o intente](https://mistral.ai/)**Modelos 13B**[Sólo CPU](https://ai.meta.com/llama/)**: Cualquier modelo funciona, sólo mucho más lento (empezar con Phi-3 Mini para la velocidad)

- Mi recomendación
- Mistral 7B
- (última versión) o
- Llama 3

## 8B

Excelente calidad para la escritura técnica

### Funciona en todos los tamaños de GPU

1. Lo suficientemente rápido para su uso interactivo[Buena para seguir las instrucciones](https://huggingface.co/models)
2. Descargando modelos`"mistral 7b gguf"`
3. Los modelos se distribuyen en Hugging Face.

**Usaremos versiones cuantificadas de GGUF.**Encontrar modelos GGUF

- [Ir a](https://huggingface.co/TheBloke/Mistral-7B-Instruct-v0.2-GGUF)
- [Cara de abrazo](https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGUF)
- [Buscar:](https://huggingface.co/QuantFactory/Meta-Llama-3-8B-Instruct-GGUF)

### Busque las cuantificaciones de TheBloke (más populares)

```bash
# Install huggingface-cli
pip install huggingface-hub

# Download Mistral 7B Q5_K_M (recommended)
huggingface-cli download TheBloke/Mistral-7B-Instruct-v0.2-GGUF \
    mistral-7b-instruct-v0.2.Q5_K_M.gguf \
    --local-dir C:\models\mistral-7b \
    --local-dir-use-symlinks False
```

Vínculos directos

1. (Las cuantificaciones de TheBloke):
2. Mistral-7B-Instruct-v0.2-GGUF`mistral-7b-instruct-v0.2.Q5_K_M.gguf`Llama-2-7B-Chat-GGUF
3. Llama-3-8B-Instruct-GGUF
4. Descargar Cuantificación específica`C:\models\mistral-7b\`

## O descargar manualmente:

### Haga clic en la pestaña "Archivos y versiones"

```bash
cd Mostlylucid.BlogLLM.Core
dotnet add package LLamaSharp  # Latest version
dotnet add package LLamaSharp.Backend.Cuda12  # Latest, matching CUDA version
```

**Buscar**

- `[LLamaSharp](https://github.com/SciSharp/LLamaSharp)`(~4,8 GB)
- `LLamaSharp.Backend.Cuda12` - [Haga clic en descargar](https://developer.nvidia.com/cuda-toolkit)Guardar en

### Configuración de LLamaSharp

Instalar el paquete NuGet

```csharp
using LLama;
using LLama.Common;

// Check if CUDA is available
bool cudaAvailable = NativeLibraryConfig.Instance.CudaEnabled;
Console.WriteLine($"CUDA Available: {cudaAvailable}");
```

¿Por qué dos paquetes?`false`- Biblioteca básica

1. CUDA
2. `LLamaSharp.Backend.Cuda12`12 binarios para la aceleración de la GPU
3. Verificar el motor CUDA

## LLamaSharp detectará automáticamente CUDA si se instala correctamente.

### Si

```csharp
using LLama;
using LLama.Common;

namespace Mostlylucid.BlogLLM.Core.Services
{
    public class ModelParameters
    {
        public string ModelPath { get; set; } = string.Empty;
        public int ContextSize { get; set; } = 4096;  // Context window
        public int GpuLayerCount { get; set; } = 35;  // Layers on GPU (35 = all for 7B)
        public int Seed { get; set; } = 1337;  // For reproducibility
        public float Temperature { get; set; } = 0.7f;  // Creativity (0.0 = deterministic, 1.0 = creative)
        public float TopP { get; set; } = 0.9f;  // Nucleus sampling
        public int MaxTokens { get; set; } = 500;  // Max generation length
    }
}
```

**, comprobar:**:

- **CUDA 12.x instalada (Parte 2)**paquete instalado
  
  - PATH incluye directorio de bin CUDA
  - Construcción del Servicio LLM

- **Parámetros del modelo**Explicaciones del parámetro
  
  - ContextSize
  - : Cuánto texto el modelo puede "ver" a la vez
  - 4096 tokens  siguientes 3000 palabras

- **Más grande = más contexto pero más lento y más VRAM**GpuLayerCount
  
  - : Cuantas capas de transformadores funcionan en la GPU
  - Los modelos 7B tienen ~32 capas
  - 35 = poner todo en la GPU (más rápido)

- **Valores inferiores = utilizar menos VRAM pero más lento**Temperatura
  
  - : Controla la aleatoriedad
  - 0,0 = elegir siempre el token más probable (aburrido, repetitivo)

### 0.7 = buen balance (por defecto)

```csharp
using LLama;
using LLama.Common;
using Microsoft.Extensions.Logging;

namespace Mostlylucid.BlogLLM.Core.Services
{
    public interface ILlmService
    {
        Task<string> GenerateAsync(string prompt, CancellationToken cancellationToken = default);
        Task<string> GenerateWithContextAsync(string prompt, List<string> contextChunks, CancellationToken cancellationToken = default);
    }

    public class LlmService : ILlmService, IDisposable
    {
        private readonly LLamaWeights _model;
        private readonly LLamaContext _context;
        private readonly ILogger<LlmService> _logger;
        private readonly ModelParameters _parameters;

        public LlmService(ModelParameters parameters, ILogger<LlmService> logger)
        {
            _parameters = parameters;
            _logger = logger;

            _logger.LogInformation("Loading model from {ModelPath}", parameters.ModelPath);

            // Configure model parameters
            var modelParams = new ModelParams(parameters.ModelPath)
            {
                ContextSize = (uint)parameters.ContextSize,
                GpuLayerCount = parameters.GpuLayerCount,
                Seed = (uint)parameters.Seed,
                UseMemoryLock = true,  // Keep model in RAM
                UseMemorymap = true    // Memory-map the model file
            };

            // Load model
            _model = LLamaWeights.LoadFromFile(modelParams);
            _context = _model.CreateContext(modelParams);

            _logger.LogInformation("Model loaded successfully. VRAM used: ~{VRAM}GB",
                EstimateVRAMUsage(parameters.GpuLayerCount));
        }

        public async Task<string> GenerateAsync(string prompt, CancellationToken cancellationToken = default)
        {
            var executor = new InteractiveExecutor(_context);

            var inferenceParams = new InferenceParams
            {
                Temperature = _parameters.Temperature,
                TopP = _parameters.TopP,
                MaxTokens = _parameters.MaxTokens,
                AntiPrompts = new[] { "\n\nUser:", "###" }  // Stop generation at these
            };

            var result = new StringBuilder();

            _logger.LogInformation("Generating response for prompt: {Prompt}", TruncateForLog(prompt));

            await foreach (var token in executor.InferAsync(prompt, inferenceParams, cancellationToken))
            {
                result.Append(token);
            }

            var response = result.ToString().Trim();
            _logger.LogInformation("Generated {Tokens} tokens", CountTokens(response));

            return response;
        }

        public async Task<string> GenerateWithContextAsync(
            string prompt,
            List<string> contextChunks,
            CancellationToken cancellationToken = default)
        {
            // Build prompt with retrieved context
            var fullPrompt = BuildContextualPrompt(prompt, contextChunks);

            _logger.LogInformation("Context chunks: {Count}, Total prompt tokens: ~{Tokens}",
                contextChunks.Count, CountTokens(fullPrompt));

            return await GenerateAsync(fullPrompt, cancellationToken);
        }

        private string BuildContextualPrompt(string userPrompt, List<string> contextChunks)
        {
            var sb = new StringBuilder();

            sb.AppendLine("You are a helpful writing assistant for a technical blog.");
            sb.AppendLine("Use the following excerpts from past blog posts as context:");
            sb.AppendLine();

            for (int i = 0; i < contextChunks.Count; i++)
            {
                sb.AppendLine($"--- Context {i + 1} ---");
                sb.AppendLine(contextChunks[i]);
                sb.AppendLine();
            }

            sb.AppendLine("---");
            sb.AppendLine();
            sb.AppendLine("Based on the context above, help with the following:");
            sb.AppendLine(userPrompt);
            sb.AppendLine();
            sb.AppendLine("Response:");

            return sb.ToString();
        }

        private int CountTokens(string text)
        {
            // Rough estimate: 1 token ≈ 4 characters
            return text.Length / 4;
        }

        private string TruncateForLog(string text, int maxLength = 100)
        {
            if (text.Length <= maxLength) return text;
            return text.Substring(0, maxLength) + "...";
        }

        private double EstimateVRAMUsage(int gpuLayers)
        {
            // Rough estimate for 7B model
            return (gpuLayers / 35.0) * 6.0;  // ~6GB for full 7B model
        }

        public void Dispose()
        {
            _context?.Dispose();
            _model?.Dispose();
        }
    }
}
```

**1.0+ = muy creativo (puede ser absurdo)**:

1. **TopP**: Muestreo de nucleo
2. **0.9 = considerar tokens que constituyen el 90% de la masa de probabilidad**Evita la toma de muestras a partir de fichas muy improbables
3. **Aplicación de los servicios de LLM**Cómo funciona
4. **Carga de modelos**: Carga el modelo GGUF en VRAM usando parámetros especificados
5. **Ejecutador Interactivo**: Modo de ejecución de LLamaSharp para interacciones similares al chat

### InferAsync

```csharp
using Microsoft.Extensions.Logging;

class Program
{
    static async Task Main(string[] args)
    {
        // Setup logging
        var loggerFactory = LoggerFactory.Create(builder => builder.AddConsole());
        var logger = loggerFactory.CreateLogger<LlmService>();

        // Configure model
        var parameters = new ModelParameters
        {
            ModelPath = @"C:\models\mistral-7b\mistral-7b-instruct-v0.2.Q5_K_M.gguf",
            ContextSize = 4096,
            GpuLayerCount = 35,
            Temperature = 0.7f,
            MaxTokens = 200
        };

        // Create service
        using var llmService = new LlmService(parameters, logger);

        // Test simple generation
        Console.WriteLine("=== Test 1: Simple Generation ===\n");
        var response1 = await llmService.GenerateAsync(
            "Explain what Docker Compose is in 2-3 sentences."
        );
        Console.WriteLine(response1);
        Console.WriteLine("\n");

        // Test with context
        Console.WriteLine("=== Test 2: Generation with Context ===\n");
        var context = new List<string>
        {
            "Docker Compose is a tool for defining and running multi-container Docker applications. With Compose, you use a YAML file to configure your application's services.",
            "In development, Docker Compose makes it easy to spin up all dependencies (databases, caches, etc.) with one command: docker-compose up."
        };

        var response2 = await llmService.GenerateWithContextAsync(
            "Write an introduction paragraph for a blog post about using Docker Compose for development dependencies.",
            context
        );
        Console.WriteLine(response2);
    }
}
```

**: Transmite tokens a medida que se generan (salida en tiempo real)**:

```
=== Test 1: Simple Generation ===

Docker Compose is a tool that allows you to define and run multi-container Docker applications using a simple YAML configuration file. It simplifies the process of managing multiple containers, networking, and volumes, making it ideal for development environments.

=== Test 2: Generation with Context ===

If you've ever found yourself juggling multiple terminal windows to start databases, caches, and other services for local development, Docker Compose is about to become your new best friend. This powerful tool lets you define your entire development environment in a single YAML file and spin everything up with one command. In this post, we'll explore how to leverage Docker Compose to manage all your development dependencies, making your local setup reproducible, shareable, and incredibly easy to manage.
```

Context Building

## : Combina el aviso del usuario con trozos de blog recuperados

Antiprompts

### : Detiene la generación en ciertas cuerdas (previene divagar)

```csharp
namespace Mostlylucid.BlogLLM.Client.Services
{
    public class SuggestionService : ISuggestionService
    {
        private readonly BatchEmbeddingService _embeddingService;
        private readonly QdrantVectorStore _vectorStore;
        private readonly ILlmService _llmService;  // NEW

        public SuggestionService(
            BatchEmbeddingService embeddingService,
            QdrantVectorStore vectorStore,
            ILlmService llmService)  // NEW
        {
            _embeddingService = embeddingService;
            _vectorStore = vectorStore;
            _llmService = llmService;
        }

        public async Task<string> GenerateAiSuggestionAsync(
            string currentText,
            List<SimilarPost> context)
        {
            // Extract text from similar posts
            var contextChunks = context
                .Take(3)  // Top 3 most similar
                .Select(p => p.FullText)
                .ToList();

            // Determine what type of suggestion to generate
            var prompt = DeterminePromptType(currentText);

            // Generate suggestion
            var suggestion = await _llmService.GenerateWithContextAsync(
                prompt,
                contextChunks
            );

            return suggestion;
        }

        private string DeterminePromptType(string currentText)
        {
            // Analyze what user is writing
            var lines = currentText.Split('\n');
            var lastLine = lines.LastOrDefault(l => !string.IsNullOrWhiteSpace(l)) ?? "";

            // Is user starting a new section?
            if (lastLine.StartsWith("## "))
            {
                return "Suggest 3-5 bullet points for what this section could cover.";
            }

            // Is user writing code?
            if (lastLine.Contains("```"))
            {
                return "Suggest relevant code examples that might be useful here.";
            }

            // Is user writing an introduction?
            if (currentText.Length < 500 && currentText.Contains("## Introduction"))
            {
                return "Suggest 2-3 sentences to continue this introduction based on similar posts.";
            }

            // Default: continue current thought
            return "Suggest 1-2 sentences to continue the current paragraph in a natural way.";
        }
    }
}
```

### Probando el servicio

```csharp
public partial class SuggestionsViewModel : ViewModelBase
{
    [RelayCommand]
    private async Task RegenerateSuggestion()
    {
        IsGenerating = true;
        AiSuggestion = "Generating...";

        try
        {
            var currentText = GetCurrentEditorText();  // From messaging
            var suggestion = await _suggestionService.GenerateAiSuggestionAsync(
                currentText,
                SimilarPosts.ToList()
            );

            AiSuggestion = suggestion;
        }
        catch (Exception ex)
        {
            AiSuggestion = $"Error: {ex.Message}";
        }
        finally
        {
            IsGenerating = false;
        }
    }
}
```

## Producto previsto

### ¡Increíble!

El modelo está trabajando y generando un texto coherente y consciente del contexto.

```csharp
public class LlmServiceFactory
{
    private static LlmService? _instance;
    private static readonly object _lock = new();

    public static LlmService GetInstance(ModelParameters parameters, ILogger<LlmService> logger)
    {
        if (_instance == null)
        {
            lock (_lock)
            {
                if (_instance == null)
                {
                    _instance = new LlmService(parameters, logger);
                }
            }
        }

        return _instance;
    }
}
```

### Integración con el Grupo de Sugerencias

Ahora vamos a integrar la generación LLM en nuestro cliente de Windows de la Parte 5.

```csharp
public class StatefulLlmService
{
    private readonly InferenceParams _defaultParams;
    private string _cachedPromptPrefix = string.Empty;

    public async Task<string> GenerateWithPrefixAsync(string prefix, string newPrompt)
    {
        // If prefix matches cached, reuse KV cache
        if (prefix == _cachedPromptPrefix)
        {
            // Only process new tokens
            return await GenerateAsync(newPrompt);
        }

        // Process entire prompt and cache
        _cachedPromptPrefix = prefix;
        return await GenerateAsync(prefix + newPrompt);
    }
}
```

Actualizar el servicio de sugerencias

### Actualizar sugerenciasModelo de vista

Optimización del rendimiento

```csharp
public async Task<List<string>> GenerateBatchAsync(List<string> prompts)
{
    var results = new List<string>();

    foreach (var prompt in prompts)
    {
        // With KV cache reuse, subsequent prompts are faster
        results.Add(await GenerateAsync(prompt));
    }

    return results;
}
```

## Caché modelo

Mantenga el modelo cargado entre solicitudes:

### Reutilización de caché KV

```csharp
private string PromptContinueWriting(string currentText, List<string> context)
{
    return $@"You are a technical blog writing assistant.

Here are excerpts from similar blog posts:
{string.Join("\n\n", context.Select((c, i) => $"--- Post {i + 1} ---\n{c}"))}

Current draft:
{currentText}

Task: Suggest 2-3 sentences to naturally continue the current paragraph.
Keep the same technical depth and casual, pragmatic tone.

Suggestion:";
}
```

### LLamaSharp admite la reutilización de la caché KV para generaciones posteriores más rápidas:

```csharp
private string PromptSectionStructure(string sectionTitle, List<string> context)
{
    return $@"You are a technical blog writing assistant.

Similar sections from past posts:
{string.Join("\n\n", context)}

New section: {sectionTitle}

Task: Suggest 4-6 bullet points for what this section should cover.
Format as a markdown list.

Bullets:";
}
```

### Esto es particularmente útil para nuestro caso de uso - los trozos de contexto permanecen igual, sólo cambia la pregunta del usuario.

```csharp
private string PromptCodeExample(string description, List<string> context)
{
    return $@"You are a C# coding assistant.

Relevant code from past posts:
{string.Join("\n\n", context)}

Task: {description}

Provide a clean, well-commented C# code example.

Code:";
}
```

## Batching

Para múltiples sugerencias, por lotes:

### Ingeniería rápida para la asistencia de escritura

```csharp
public class VramMonitor
{
    [DllImport("nvml.dll")]
    private static extern int nvmlDeviceGetMemoryInfo(IntPtr device, ref NvmlMemory memory);

    [StructLayout(LayoutKind.Sequential)]
    public struct NvmlMemory
    {
        public ulong Total;
        public ulong Free;
        public ulong Used;
    }

    public static (ulong used, ulong total) GetVramUsage()
    {
        // Simplified - actual implementation needs proper NVML initialization
        var memory = new NvmlMemory();
        // nvmlDeviceGetMemoryInfo(device, ref memory);

        return (memory.Used / 1024 / 1024, memory.Total / 1024 / 1024);  // Convert to MB
    }
}
```

### Buenas indicaciones = buena salida.

```csharp
public class LlmServiceWithUnload : IDisposable
{
    private LlmService? _service;
    private readonly Timer _unloadTimer;
    private DateTime _lastUsed;

    public LlmServiceWithUnload()
    {
        _unloadTimer = new Timer(CheckForUnload, null, TimeSpan.FromMinutes(1), TimeSpan.FromMinutes(1));
    }

    private void CheckForUnload(object? state)
    {
        if (_service != null && (DateTime.Now - _lastUsed) > TimeSpan.FromMinutes(10))
        {
            _service.Dispose();
            _service = null;
            GC.Collect();
            Console.WriteLine("Model unloaded due to inactivity");
        }
    }

    public async Task<string> GenerateAsync(string prompt)
    {
        _lastUsed = DateTime.Now;

        if (_service == null)
        {
            // Reload model
            _service = CreateService();
        }

        return await _service.GenerateAsync(prompt);
    }
}
```

## Aquí están las plantillas para diferentes escenarios:

Continuar escribiendo

```csharp
public async Task<string> GenerateWithRetryAsync(string prompt, int maxRetries = 3)
{
    for (int i = 0; i < maxRetries; i++)
    {
        try
        {
            return await GenerateAsync(prompt);
        }
        catch (OutOfMemoryException)
        {
            _logger.LogWarning("OOM error, reducing max tokens");
            _parameters.MaxTokens = Math.Max(100, _parameters.MaxTokens / 2);
        }
        catch (Exception ex)
        {
            _logger.LogError(ex, "Generation failed, attempt {Attempt}/{Max}", i + 1, maxRetries);

            if (i == maxRetries - 1) throw;

            await Task.Delay(1000 * (i + 1));  // Exponential backoff
        }
    }

    throw new Exception("Generation failed after retries");
}
```

## Sugerir estructura de sección

Ejemplo de código

1. ✅ Chose [Gestión de memoria](https://github.com/SciSharp/LLamaSharp)Con grandes modelos, la gestión de la memoria es crucial.
2. ✅ Understood [Monitor de uso de VRAM](https://github.com/ggerganov/ggml/blob/master/docs/gguf.md)Descarga el modelo cuando esté inactivo
3. ✅ Selected appropriate model ([Manejo de errores](https://mistral.ai/) / [Los LLM pueden fallar de maneras inesperadas.](https://ai.meta.com/llama/)Manejar con gracia:
4. ✅ Implemented LlmService with CUDA acceleration
5. ✅ Integrated with Windows client for suggestions
6. ✅ Implemented prompt engineering for writing tasks
7. ✅ Added performance optimizations (caching, batching)
8. ✅ Handled memory management and errors

## Resumen

Hemos integrado con éxito la inferencia local de LLM:**[LLAMASharp](/blog/building-a-lawyer-gpt-for-your-blog-part7)**para la integración C#

- Formato GGUF
- y cuantificación
- Mistral 7B
- Llama 3
- 8B)
- ¿Qué sigue?
- In

Parte 7: Generación de contenido e ingeniería rápida

## , nos centraremos en la tubería completa de generación de contenido:

- [Técnicas avanzadas de ingeniería rápida](/blog/building-a-lawyer-gpt-for-your-blog-part1)
- [Conversación multiturno para el refinamiento iterativo](/blog/building-a-lawyer-gpt-for-your-blog-part2)
- [Estrategias de gestión de ventanas de contexto](/blog/building-a-lawyer-gpt-for-your-blog-part3)
- [Evaluación y filtrado de la calidad](/blog/building-a-lawyer-gpt-for-your-blog-part4)
- [Cumplimiento de la coherencia del estilo](/blog/building-a-lawyer-gpt-for-your-blog-part5)
- **Gestión de la generación de bloques de código**Patrones de uso en el mundo real
- [¡Vamos a hacer que el sistema realmente útil para la escritura diaria del blog!](/blog/building-a-lawyer-gpt-for-your-blog-part7)
- [Navegación en serie](/blog/building-a-lawyer-gpt-for-your-blog-part8)

## Parte 1: Introducción y Arquitectura

- [Parte 2: Configuración de la GPU y CUDA en C#](https://scisharp.github.io/LLamaSharp/)
- [Parte 3: Comprender los embebidos y las bases de datos vectoriales](https://github.com/ggerganov/ggml/blob/master/docs/gguf.md)
- [Parte 4: Construcción del oleoducto de ingestión](https://huggingface.co/TheBloke)
- [Parte 5: El cliente de Windows](https://github.com/ggerganov/llama.cpp)

Parte 6: Integración local de la LLM[(en este puesto)](/blog/building-a-lawyer-gpt-for-your-blog-part7)!