Safetensors
llama
航空航天
中英文语言模型

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

You agree to not use the dataset to conduct experiments that cause harm to human subjects.

Log in or Sign Up to review the conditions and access this model content.

This model is finetuned on the model llama3.1-8b-instruct using the dataset BAAI/IndustryInstruction_Aerospace dataset, the dataset details can jump to the repo: BAAI/IndustryInstruction

training params

The training framework is llama-factory, template=llama3

learning_rate=1e-5
lr_scheduler_type=cosine
max_length=2048
warmup_ratio=0.05
batch_size=64
epoch=10

select best ckpt by the evaluation loss

evaluation

Duto to there is no evaluation benchmark, we can not eval the model

How to use

# !/usr/bin/env python
# -*- coding:utf-8 -*-
# ==================================================================
# [Author]       : xiaofeng
# [Descriptions] :
# ==================================================================

from transformers import AutoTokenizer, AutoModelForCausalLM
import transformers
import torch


llama3_jinja = """{% if messages[0]['role'] == 'system' %}
    {% set offset = 1 %}
{% else %}
    {% set offset = 0 %}
{% endif %}

{{ bos_token }}
{% for message in messages %}
    {% if (message['role'] == 'user') != (loop.index0 % 2 == offset) %}
        {{ raise_exception('Conversation roles must alternate user/assistant/user/assistant/...') }}
    {% endif %}

    {{ '<|start_header_id|>' + message['role'] + '<|end_header_id|>\n\n' + message['content'] | trim + '<|eot_id|>' }}
{% endfor %}

{% if add_generation_prompt %}
    {{ '<|start_header_id|>' + 'assistant' + '<|end_header_id|>\n\n' }}
{% endif %}"""


dtype = torch.bfloat16

model_dir = "XiaofengAlg/Aerospace-llama3_1_8B_instruct"
model = AutoModelForCausalLM.from_pretrained(
    model_dir,
    device_map="cuda",
    torch_dtype=dtype,
)

tokenizer = AutoTokenizer.from_pretrained(model_dir)
tokenizer.chat_template = llama3_jinja  # update template

message = [
    {"role": "system", "content": "You are a helpful assistant"},
    {
        "role": "user",
        "content": "神舟十六号在中国空间站中的角色是什么,它如何实现对太空的全面覆盖?",
    },
]
prompt = tokenizer.apply_chat_template(
    message, tokenize=False, add_generation_prompt=True
)
print(prompt)
inputs = tokenizer.encode(prompt, add_special_tokens=False, return_tensors="pt")
prompt_length = len(inputs[0])
print(f"prompt_length:{prompt_length}")

generating_args = {
    "do_sample": True,
    "temperature": 1.0,
    "top_p": 0.5,
    "top_k": 15,
    "max_new_tokens": 512,
}


generate_output = model.generate(input_ids=inputs.to(model.device), **generating_args)

response_ids = generate_output[:, prompt_length:]
response = tokenizer.batch_decode(
    response_ids, skip_special_tokens=True, clean_up_tokenization_spaces=True
)[0]


"""
神舟十六号作为中国空间站的最后一批次载人飞船,它将在空间站的建造和运营中发挥关键作用。神舟十六号将与天舟四号货运飞船和天舟五号货运飞船对接,共同完成空间站的组装,并进行空间站的全面覆盖。神舟十六号的发射将标志着中国空间站的全面建成,并且将为中国航天员在太空中进行长期驻留和科学实验提供支持。
"""
print(f"response:{response}")

Citation

@misc{shi2024industryinstruction,
  title     = {IndustryInstruction},
  author    = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao and Yonghua Lin},
  year      = {2024},
  publisher = {Hugging Face},
  doi       = {10.57967/hf/3487},
  url       = {https://huggingface.co/datasets/BAAI/IndustryInstruction}
}
Downloads last month
15
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Model tree for XiaofengAlg/Aerospace-llama3_1_8B_instruct

Finetuned
(3255)
this model

Datasets used to train XiaofengAlg/Aerospace-llama3_1_8B_instruct

Collection including XiaofengAlg/Aerospace-llama3_1_8B_instruct