This model is finetuned on the model llama3.1-8b-instruct using the dataset BAAI/IndustryInstruction_Aerospace dataset, the dataset details can jump to the repo: BAAI/IndustryInstruction
training params
The training framework is llama-factory, template=llama3
learning_rate=1e-5
lr_scheduler_type=cosine
max_length=2048
warmup_ratio=0.05
batch_size=64
epoch=10
select best ckpt by the evaluation loss
evaluation
Duto to there is no evaluation benchmark, we can not eval the model
How to use
# !/usr/bin/env python
# -*- coding:utf-8 -*-
# ==================================================================
# [Author] : xiaofeng
# [Descriptions] :
# ==================================================================
from transformers import AutoTokenizer, AutoModelForCausalLM
import transformers
import torch
llama3_jinja = """{% if messages[0]['role'] == 'system' %}
{% set offset = 1 %}
{% else %}
{% set offset = 0 %}
{% endif %}
{{ bos_token }}
{% for message in messages %}
{% if (message['role'] == 'user') != (loop.index0 % 2 == offset) %}
{{ raise_exception('Conversation roles must alternate user/assistant/user/assistant/...') }}
{% endif %}
{{ '<|start_header_id|>' + message['role'] + '<|end_header_id|>\n\n' + message['content'] | trim + '<|eot_id|>' }}
{% endfor %}
{% if add_generation_prompt %}
{{ '<|start_header_id|>' + 'assistant' + '<|end_header_id|>\n\n' }}
{% endif %}"""
dtype = torch.bfloat16
model_dir = "MonteXiaofeng/Aerospace-llama3_1_8B_instruct"
model = AutoModelForCausalLM.from_pretrained(
model_dir,
device_map="cuda",
torch_dtype=dtype,
)
tokenizer = AutoTokenizer.from_pretrained(model_dir)
tokenizer.chat_template = llama3_jinja # update template
message = [
{"role": "system", "content": "You are a helpful assistant"},
{
"role": "user",
"content": "神舟十六号在中国空间站中的角色是什么,它如何实现对太空的全面覆盖?",
},
]
prompt = tokenizer.apply_chat_template(
message, tokenize=False, add_generation_prompt=True
)
print(prompt)
inputs = tokenizer.encode(prompt, add_special_tokens=False, return_tensors="pt")
prompt_length = len(inputs[0])
print(f"prompt_length:{prompt_length}")
generating_args = {
"do_sample": True,
"temperature": 1.0,
"top_p": 0.5,
"top_k": 15,
"max_new_tokens": 512,
}
generate_output = model.generate(input_ids=inputs.to(model.device), **generating_args)
response_ids = generate_output[:, prompt_length:]
response = tokenizer.batch_decode(
response_ids, skip_special_tokens=True, clean_up_tokenization_spaces=True
)[0]
"""
神舟十六号作为中国空间站的最后一批次载人飞船,它将在空间站的建造和运营中发挥关键作用。神舟十六号将与天舟四号货运飞船和天舟五号货运飞船对接,共同完成空间站的组装,并进行空间站的全面覆盖。神舟十六号的发射将标志着中国空间站的全面建成,并且将为中国航天员在太空中进行长期驻留和科学实验提供支持。
"""
print(f"response:{response}")
- Downloads last month
- 6
Model tree for MonteXiaofeng/Aerospace-llama3_1_8B_instruct
Base model
meta-llama/Llama-3.1-8B
Finetuned
meta-llama/Llama-3.1-8B-Instruct