Text Generation
Transformers
Safetensors
PEFT
English
text-generation-inference
unsloth
qwen2
qlora
sql
text-to-sql
code-generation
conversational
Eval Results (legacy)
Instructions to use DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora", device_map="auto") - PEFT
How to use DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora
- SGLang
How to use DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Studio
How to use DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora", max_seq_length=2048, ) - Docker Model Runner
How to use DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora with Docker Model Runner:
docker model run hf.co/DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora
Add comprehensive model card with training details, usage examples, and limitations
5668a9a verified | base_model: unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit | |
| library_name: transformers | |
| pipeline_tag: text-generation | |
| license: apache-2.0 | |
| language: | |
| - en | |
| tags: | |
| - text-generation-inference | |
| - transformers | |
| - unsloth | |
| - qwen2 | |
| - qlora | |
| - peft | |
| - sql | |
| - text-to-sql | |
| - code-generation | |
| datasets: | |
| - DanielRegaladoCardoso/text-to-sql-mix-v2 | |
| model-index: | |
| - name: sql-generator-qwen25-coder-7b-lora | |
| results: | |
| - task: | |
| type: text-generation | |
| name: Text-to-SQL Generation | |
| dataset: | |
| name: text-to-sql-mix-v2 | |
| type: DanielRegaladoCardoso/text-to-sql-mix-v2 | |
| metrics: | |
| - type: train_loss | |
| value: 0.2658 | |
| name: Final training loss | |
| # SQL Generator β Qwen2.5-Coder-7B (QLoRA) | |
| Fine-tuned **Qwen2.5-Coder-7B-Instruct** for **text-to-SQL** generation. Given a SQL schema and a natural-language question, the model produces a syntactically correct SQL query. | |
| Trained as part of the [SQL Agent LLMOps](https://github.com/DanielRegaladoUMiami/sql-agent-llmops) project β a multi-model SQL agent with deployment on HuggingFace Spaces. | |
| ## Model details | |
| | | | | |
| |---|---| | |
| | **Base model** | [unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit](https://huggingface.co/unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit) | | |
| | **Architecture** | Qwen2 (7.6B params, 4-bit quantized base) | | |
| | **Fine-tuning method** | QLoRA via [Unsloth](https://github.com/unslothai/unsloth) + TRL | | |
| | **LoRA rank** | 16 | | |
| | **LoRA alpha** | 32 | | |
| | **Target modules** | `q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj` | | |
| | **Trainable params** | ~70 M (~0.9% of base) | | |
| | **Language** | English | | |
| | **License** | Apache 2.0 | | |
| ## Training data | |
| [`DanielRegaladoCardoso/text-to-sql-mix-v2`](https://huggingface.co/datasets/DanielRegaladoCardoso/text-to-sql-mix-v2) β a curated mix of 5 public text-to-SQL datasets: | |
| - `b-mc2/sql-create-context` | |
| - `gretelai/synthetic_text_to_sql` | |
| - `knowrohit07/know_sql` | |
| - `NumbersStation/NSText2SQL` | |
| - `Clinton/Text-to-sql-v1` | |
| **Final training set**: 672,949 examples (after filtering sequences > 1024 tokens β kept 93.1% of original 723,097 rows). With sequence packing, this compressed to 154,462 effective sequences of length 1024. | |
| ## Training configuration | |
| | Hyperparameter | Value | | |
| |---|---| | |
| | Hardware | 1Γ NVIDIA L40S (48 GB) | | |
| | Epochs | 1 | | |
| | Batch size (per device) | 16 | | |
| | Gradient accumulation | 1 | | |
| | Effective batch size | 16 | | |
| | Max sequence length | 1024 | | |
| | Learning rate | 1e-4 | | |
| | LR scheduler | Cosine | | |
| | Warmup ratio | 0.03 | | |
| | Optimizer | adamw_8bit | | |
| | Precision | bf16 | | |
| | Sequence packing | Enabled | | |
| | Total steps | 9,654 | | |
| | Wall-clock time | 13.5 hours | | |
| | Final training loss | **0.2658** | | |
| ## Prompt format | |
| The model expects a chat-style prompt with a system message defining the SQL-expert role and a user message containing the schema and question: | |
| ``` | |
| <|im_start|>system | |
| You are a SQL expert. Given a SQL schema and a natural-language question, generate a correct SQL query answering the question. Return only the SQL. | |
| <|im_end|> | |
| <|im_start|>user | |
| ### Schema | |
| CREATE TABLE players (id INT, name VARCHAR, hometown VARCHAR); | |
| ### Question | |
| List all players from Tampa, Florida. | |
| <|im_end|> | |
| <|im_start|>assistant | |
| ``` | |
| ## Usage | |
| ### Option A β Load merged 16-bit model (recommended) | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| REPO = "DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora" | |
| model = AutoModelForCausalLM.from_pretrained(REPO, torch_dtype="auto", device_map="auto") | |
| tokenizer = AutoTokenizer.from_pretrained(REPO) | |
| messages = [ | |
| {"role": "system", "content": "You are a SQL expert. Given a SQL schema and a natural-language question, generate a correct SQL query answering the question. Return only the SQL."}, | |
| {"role": "user", "content": "### Schema\nCREATE TABLE players (id INT, name VARCHAR, hometown VARCHAR);\n\n### Question\nList all players from Tampa, Florida."}, | |
| ] | |
| input_ids = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt").to(model.device) | |
| out = model.generate(input_ids, max_new_tokens=256, do_sample=False) | |
| print(tokenizer.decode(out[0][input_ids.shape[1]:], skip_special_tokens=True)) | |
| # β SELECT * FROM players WHERE hometown = 'Tampa, Florida' | |
| ``` | |
| ### Option B β Load LoRA adapter on top of base model | |
| Useful if you want to keep the base model in 4-bit (lower VRAM footprint). | |
| ```python | |
| from peft import PeftModel | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| import torch | |
| base = AutoModelForCausalLM.from_pretrained( | |
| "Qwen/Qwen2.5-Coder-7B-Instruct", | |
| torch_dtype=torch.bfloat16, | |
| device_map="auto", | |
| ) | |
| model = PeftModel.from_pretrained(base, "DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora") | |
| tokenizer = AutoTokenizer.from_pretrained("DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora") | |
| ``` | |
| ### Option C β Inference with Unsloth (fastest) | |
| ```python | |
| from unsloth import FastLanguageModel | |
| model, tokenizer = FastLanguageModel.from_pretrained( | |
| "DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora", | |
| max_seq_length=1024, | |
| load_in_4bit=True, | |
| ) | |
| FastLanguageModel.for_inference(model) | |
| ``` | |
| ## Files | |
| | File | Size | Purpose | | |
| |---|---|---| | |
| | `adapter_model.safetensors` | 161 MB | LoRA adapter weights | | |
| | `adapter_config.json` | 1 KB | LoRA configuration | | |
| | `model-0000{1..4}-of-00004.safetensors` | 15.2 GB | Merged 16-bit model | | |
| | `tokenizer.json` + `tokenizer_config.json` | 11 MB | Tokenizer | | |
| | `chat_template.jinja` | 4 KB | Qwen chat template | | |
| ## Limitations | |
| - **English only** β training data is English; performance on other languages is not validated. | |
| - **Sequence length cap** β examples requiring > 1024 tokens (large schemas, complex multi-CTE queries) were filtered out during training. The model may underperform on inputs above this length. | |
| - **No execution validation** β the model is trained to produce syntactically correct SQL, but generated queries are not guaranteed to execute or return correct results without manual review. Always sanity-check against your real database. | |
| - **Single dialect bias** β training data mixes multiple SQL dialects (SQLite, ANSI, MySQL); the model may produce queries that lean toward one dialect over another. | |
| ## Citation | |
| If you use this model, please cite the [SQL Agent LLMOps project](https://github.com/DanielRegaladoUMiami/sql-agent-llmops). | |
| ```bibtex | |
| @misc{regalado2026sqlagent, | |
| author = {Daniel Regalado Cardoso}, | |
| title = {SQL Generator: Qwen2.5-Coder-7B fine-tuned for text-to-SQL}, | |
| year = {2026}, | |
| howpublished = {\url{https://huggingface.co/DanielRegaladoCardoso/sql-generator-qwen25-coder-7b-lora}}, | |
| } | |
| ``` | |
| ## Acknowledgments | |
| - [Unsloth](https://github.com/unslothai/unsloth) β 2Γ faster QLoRA training | |
| - [TRL](https://github.com/huggingface/trl) β SFTTrainer | |
| - Qwen team β Qwen2.5-Coder-7B base model | |
| - All authors of the source datasets cited above | |
| [<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>](https://github.com/unslothai/unsloth) | |