Instructions to use Davis426/COMP8420-Healthcare-LLM-Assistant with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries
PEFT
How to use Davis426/COMP8420-Healthcare-LLM-Assistant with PEFT:
```
Task type is invalid.
```

How to use Davis426/COMP8420-Healthcare-LLM-Assistant with llama-cpp-python:

# !pip install llama-cpp-python

from llama_cpp import Llama

llm = Llama.from_pretrained(
	repo_id="Davis426/COMP8420-Healthcare-LLM-Assistant",
	filename="llama32/llama32-medqa-gguf/model.Q4_K_M.gguf",
)

llm.create_chat_completion(
	messages = [
		{
			"role": "user",
			"content": "What is the capital of France?"
		}
	]
)

Notebooks
Google Colab
Kaggle
Local Apps

llama.cpp

How to use Davis426/COMP8420-Healthcare-LLM-Assistant with llama.cpp:

Install from brew

brew install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama-server -hf Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M
# Run inference directly in the terminal:
llama-cli -hf Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M

Install from WinGet (Windows)

winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama-server -hf Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M
# Run inference directly in the terminal:
llama-cli -hf Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M

Use pre-built binary

# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M

Build from source code

git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M

Use Docker

docker model run hf.co/Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M

LM Studio
Jan

vLLM

How to use Davis426/COMP8420-Healthcare-LLM-Assistant with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "Davis426/COMP8420-Healthcare-LLM-Assistant"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "Davis426/COMP8420-Healthcare-LLM-Assistant",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker

docker model run hf.co/Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M

Ollama
How to use Davis426/COMP8420-Healthcare-LLM-Assistant with Ollama:
```
ollama run hf.co/Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M
```

Unsloth Studio new

How to use Davis426/COMP8420-Healthcare-LLM-Assistant with Unsloth Studio:

Install Unsloth Studio (macOS, Linux, WSL)

curl -fsSL https://unsloth.ai/install.sh | sh
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for Davis426/COMP8420-Healthcare-LLM-Assistant to start chatting

Install Unsloth Studio (Windows)

irm https://unsloth.ai/install.ps1 | iex
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for Davis426/COMP8420-Healthcare-LLM-Assistant to start chatting

Using HuggingFace Spaces for Unsloth

# No setup required
# Open https://huggingface.co/spaces/unsloth/studio in your browser
# Search for Davis426/COMP8420-Healthcare-LLM-Assistant to start chatting

Pi new

How to use Davis426/COMP8420-Healthcare-LLM-Assistant with Pi:

Start the llama.cpp server

# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama-server -hf Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M

Configure the model in Pi

# Install Pi:
npm install -g @mariozechner/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
  "providers": {
    "llama-cpp": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "none",
      "models": [
        {
          "id": "Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M"
        }
      ]
    }
  }
}

Run Pi

# Start Pi in your project directory:
pi

Hermes Agent new

How to use Davis426/COMP8420-Healthcare-LLM-Assistant with Hermes Agent:

Start the llama.cpp server

# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama-server -hf Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M

Configure Hermes

# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M

Run Hermes

hermes

Docker Model Runner
How to use Davis426/COMP8420-Healthcare-LLM-Assistant with Docker Model Runner:
```
docker model run hf.co/Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M
```

Lemonade

How to use Davis426/COMP8420-Healthcare-LLM-Assistant with Lemonade:

Pull the model

# Download Lemonade from https://lemonade-server.ai/
lemonade pull Davis426/COMP8420-Healthcare-LLM-Assistant:Q4_K_M

Run and chat with the model

lemonade run user.COMP8420-Healthcare-LLM-Assistant-Q4_K_M

List all available models

lemonade list

Davis426 commited on about 15 hours ago

Commit

2ec559e

verified ·

1 Parent(s): 6d281dd

Delete qwen-medqa-gguf

Browse files

Files changed (6) hide show

qwen-medqa-gguf/Modelfile +0 -13
qwen-medqa-gguf/chat_template.jinja +0 -54
qwen-medqa-gguf/config.json +0 -62
qwen-medqa-gguf/model.Q4_K_M.gguf +0 -3
qwen-medqa-gguf/tokenizer.json +0 -3
qwen-medqa-gguf/tokenizer_config.json +0 -16

qwen-medqa-gguf/Modelfile DELETED Viewed

@@ -1,13 +0,0 @@
-FROM ./model.Q4_K_M.gguf
-TEMPLATE """{{ if .System }}<|im_start|>system
-{{ .System }}<|im_end|>
-{{ end }}{{ if .Prompt }}<|im_start|>user
-{{ .Prompt }}<|im_end|>
-{{ end }}<|im_start|>assistant
-{{ .Response }}<|im_end|>"""
-PARAMETER stop "<|im_start|>"
-PARAMETER stop "<|im_end|>"
-PARAMETER temperature 0.3
-PARAMETER top_p 0.9

qwen-medqa-gguf/chat_template.jinja DELETED Viewed

@@ -1,54 +0,0 @@
-{%- if tools %}
-    {{- '<|im_start|>system\n' }}
-    {%- if messages[0]['role'] == 'system' %}
-        {{- messages[0]['content'] }}
-    {%- else %}
-        {{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
-    {%- endif %}
-    {{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
-    {%- for tool in tools %}
-        {{- "\n" }}
-        {{- tool | tojson }}
-    {%- endfor %}
-    {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
-{%- else %}
-    {%- if messages[0]['role'] == 'system' %}
-        {{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
-    {%- else %}
-        {{- '<|im_start|>system\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n' }}
-    {%- endif %}
-{%- endif %}
-{%- for message in messages %}
-    {%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
-        {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
-    {%- elif message.role == "assistant" %}
-        {{- '<|im_start|>' + message.role }}
-        {%- if message.content %}
-            {{- '\n' + message.content }}
-        {%- endif %}
-        {%- for tool_call in message.tool_calls %}
-            {%- if tool_call.function is defined %}
-                {%- set tool_call = tool_call.function %}
-            {%- endif %}
-            {{- '\n<tool_call>\n{"name": "' }}
-            {{- tool_call.name }}
-            {{- '", "arguments": ' }}
-            {{- tool_call.arguments | tojson }}
-            {{- '}\n</tool_call>' }}
-        {%- endfor %}
-        {{- '<|im_end|>\n' }}
-    {%- elif message.role == "tool" %}
-        {%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
-            {{- '<|im_start|>user' }}
-        {%- endif %}
-        {{- '\n<tool_response>\n' }}
-        {{- message.content }}
-        {{- '\n</tool_response>' }}
-        {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
-            {{- '<|im_end|>\n' }}
-        {%- endif %}
-    {%- endif %}
-{%- endfor %}
-{%- if add_generation_prompt %}
-    {{- '<|im_start|>assistant\n' }}
-{%- endif %}

qwen-medqa-gguf/config.json DELETED Viewed

@@ -1,62 +0,0 @@
-{
-    "architectures": [
-        "Qwen2ForCausalLM"
-    ],
-    "attention_dropout": 0.0,
-    "bos_token_id": null,
-    "torch_dtype": "bfloat16",
-    "eos_token_id": 151645,
-    "hidden_act": "silu",
-    "hidden_size": 1536,
-    "initializer_range": 0.02,
-    "intermediate_size": 8960,
-    "layer_types": [
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention",
-        "full_attention"
-    ],
-    "max_position_embeddings": 32768,
-    "max_window_layers": 21,
-    "model_type": "qwen2",
-    "num_attention_heads": 12,
-    "num_hidden_layers": 28,
-    "num_key_value_heads": 2,
-    "pad_token_id": 151665,
-    "rms_norm_eps": 1e-06,
-    "rope_parameters": {
-        "rope_theta": 1000000.0,
-        "rope_type": "default"
-    },
-    "sliding_window": null,
-    "tie_word_embeddings": true,
-    "unsloth_fixed": true,
-    "unsloth_version": "2026.3.17",
-    "use_cache": true,
-    "use_sliding_window": false,
-    "vocab_size": 151936
-}

qwen-medqa-gguf/model.Q4_K_M.gguf DELETED Viewed

@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:eafe61cc3011b9ced0f179246b2b0aafaf37437c2699920cad49aebeebcc494f
-size 986047968

qwen-medqa-gguf/tokenizer.json DELETED Viewed

@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:bd5948af71b4f56cf697f7580814c7ce8b80595ef985544efcacf716126a2e31
-size 11422356

qwen-medqa-gguf/tokenizer_config.json DELETED Viewed

@@ -1,16 +0,0 @@
-{
-  "add_prefix_space": false,
-  "backend": "tokenizers",
-  "bos_token": null,
-  "clean_up_tokenization_spaces": false,
-  "eos_token": "<|im_end|>",
-  "errors": "replace",
-  "is_local": false,
-  "model_max_length": 32768,
-  "pad_token": "<|PAD_TOKEN|>",
-  "padding_side": "right",
-  "split_special_tokens": false,
-  "tokenizer_class": "Qwen2Tokenizer",
-  "unk_token": null,
-  "chat_template": "{%- if tools %}\n    {{- '<|im_start|>system\\n' }}\n    {%- if messages[0]['role'] == 'system' %}\n        {{- messages[0]['content'] }}\n    {%- else %}\n        {{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}\n    {%- endif %}\n    {{- \"\\n\\n# Tools\\n\\nYou may call one or more functions to assist with the user query.\\n\\nYou are provided with function signatures within <tools></tools> XML tags:\\n<tools>\" }}\n    {%- for tool in tools %}\n        {{- \"\\n\" }}\n        {{- tool | tojson }}\n    {%- endfor %}\n    {{- \"\\n</tools>\\n\\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\\n<tool_call>\\n{\\\"name\\\": <function-name>, \\\"arguments\\\": <args-json-object>}\\n</tool_call><|im_end|>\\n\" }}\n{%- else %}\n    {%- if messages[0]['role'] == 'system' %}\n        {{- '<|im_start|>system\\n' + messages[0]['content'] + '<|im_end|>\\n' }}\n    {%- else %}\n        {{- '<|im_start|>system\\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\\n' }}\n    {%- endif %}\n{%- endif %}\n{%- for message in messages %}\n    {%- if (message.role == \"user\") or (message.role == \"system\" and not loop.first) or (message.role == \"assistant\" and not message.tool_calls) %}\n        {{- '<|im_start|>' + message.role + '\\n' + message.content + '<|im_end|>' + '\\n' }}\n    {%- elif message.role == \"assistant\" %}\n        {{- '<|im_start|>' + message.role }}\n        {%- if message.content %}\n            {{- '\\n' + message.content }}\n        {%- endif %}\n        {%- for tool_call in message.tool_calls %}\n            {%- if tool_call.function is defined %}\n                {%- set tool_call = tool_call.function %}\n            {%- endif %}\n            {{- '\\n<tool_call>\\n{\"name\": \"' }}\n            {{- tool_call.name }}\n            {{- '\", \"arguments\": ' }}\n            {{- tool_call.arguments | tojson }}\n            {{- '}\\n</tool_call>' }}\n        {%- endfor %}\n        {{- '<|im_end|>\\n' }}\n    {%- elif message.role == \"tool\" %}\n        {%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != \"tool\") %}\n            {{- '<|im_start|>user' }}\n        {%- endif %}\n        {{- '\\n<tool_response>\\n' }}\n        {{- message.content }}\n        {{- '\\n</tool_response>' }}\n        {%- if loop.last or (messages[loop.index0 + 1].role != \"tool\") %}\n            {{- '<|im_end|>\\n' }}\n        {%- endif %}\n    {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n    {{- '<|im_start|>assistant\\n' }}\n{%- endif %}\n"
-}