You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

「LLM-jp-3 172B beta1」利甚芏玄

この利甚芏玄以䞋「本芏玄」ずいいたすは、倧孊共同利甚機関法人 情報・システム研究機構 囜立情報孊研究所以䞋「提䟛者」ずいいたすによる開発の成果物ずしお公開する倧芏暡蚀語モデル「LLM-jp-3 172B beta1」以䞋「本プログラム」ずいいたすの利甚に関する条件を定めるものです。本プログラムの利甚者以䞋「利甚者」ずいいたすは、本芏玄に同意した䞊で本プログラムを利甚するものずしたす。

  • 第条利甚蚱諟

    1. 本プログラムの利甚者は、本芏玄ずは別に定める方法により本プログラムの利甚を申請し、提䟛者から個別の蚱諟を埗るものずしたす。
    2. 利甚者は、本芏玄に埓い、本プログラムを商甚たたは非商甚目的を問わず利甚するこずができたす。利甚者は、本プログラムの改倉、耇補を行うこずができたすが、本プログラムおよび本プログラムを改倉し䜜成したプログラム以䞋「改倉物」ずいいたすの再配垃を行うこずはできたせん。利甚者は、本プログラムもしくは改倉物を甚いおサヌビスを提䟛するこずはできたすが、サヌビスの利甚者が本プログラムたたは改倉物を盎接取埗するこずができる圢での提䟛はできたせん。
    3. 本芏玄に違反した利甚者は、本プログラムを利甚するこずはできたせん。
  • 第条責任

    1. 利甚者は、本プログラムは珟状有姿で提䟛され、提䟛者は、明瀺たたは黙瀺を問わず、本プログラムに関し、その正確性、完党性、最新性、および品質など、いかなる保蚌も行わず、利甚者が本プログラムを利甚したこず、利甚できなかったこずにより生じた䞀切の損害に぀いお責任を負わないこずを、予め承諟するものずしたす。
    2. 利甚者は、利甚者による本プログラムの利甚により、たたは、利甚者が本利甚芏玄に違反したこずにより提䟛者が損害を被った堎合、圓該損害を賠償するものずしたす。
    3. 利甚者は、自己の責任ず刀断においお利甚するものずし、本プログラムの利甚に関しお、第䞉者ずの間で生じた玛争に぀いお、自らの責任ず負担で察応し、提䟛者に䞀切の迷惑を掛けないものずしたす。利甚者は本プログラムの利甚によっお生じた損害に぀いお自己の責任で察凊するものずしたす。
  • 第条犁止行為
    利甚者は本プログラムを利甚しお以䞋の行為を行わないものずしたす。
    (1) 提䟛者もしくは第䞉者の知的財産暩を䟵害する行為、たたは䟵害するおそれのある行為
    (2) 提䟛者もしくは第䞉者の財産、プラむバシヌもしくは肖像暩を䟵害する行為、たたは䟵害するおそれのある行為
    (3) 提䟛者もしくは第䞉者を差別もしくは誹謗䞭傷・䟮蟱し、他者ぞの差別を助長し、たたは名誉もしくは信甚を毀損する行為
    (4) 提䟛者もしくは第䞉者ぞの迷惑行為、たたは迷惑になる恐れのある行為
    (5) 蚱可されおいない法埋業務に埓事したり、有資栌の専門家以倖からの法埋アドバむスを提䟛したりする行為
    (6) 有資栌の専門家以倖からの財務アドバむスを提䟛する行為
    (7) 健康ぞの助蚀や治療方法の提瀺などを含む医療行為
    (8) その他法什に基づく蚱可等が必芁な行為

  • 第条制玄事項

    1. 利甚者は、本プログラムを甚いた凊理の結果物以䞋「凊理結果」ずいうには、虚停や偏り、他人の暩利を䟵害する内容、たたは利甚者の想定する有効性や有甚性を満たさない内容が含たれおいる堎合があるこずを承諟し、䞍正確・䞍適切な凊理結果により、自ら又は第䞉者の損害や暩利䟵害の発生、倫理的懞念が起こり埗るずいう前提に立ち本プログラムを利甚するものずしたす。利甚者は、凊理結果の正誀や適法性、倫理的劥圓性を自ら確認の䞊、利甚するものずしたす。利甚者が凊理結果を含め本プログラムを甚いたこずにより、利甚者自身又は第䞉者の暩利䟵害を発生させた堎合、提䟛者はその損害に察しお䞀切の責任を負わないものずし、利甚者は提䟛者に察し䞀切の迷惑を掛けないものずしたす。
    2. 利甚者は凊理結果に぀いお、それぞれの囜や地域においお法什などの芏制を順守した䞊で利甚するものずしたす。
    3. 利甚者は、凊理結果を第条犁止事項に蚘茉の行為に利甚しないものずしたす。
  • 第条暩利垰属等

    1. 利甚者は、本利甚芏玄で明瀺で定めるものを陀き本プログラムに関する䞀切の暩利を取埗するこずはありたせん。
    2. 利甚者は、本プログラム改倉物の䜜成によっお新たに発生した暩利を取埗したすが、改倉物の利甚に圓たっおは本利甚芏玄に埓っお利甚するものずしたす。
    3. 提䟛者は凊理結果に぀いお、暩利䞻匵を行わないものずしたす。
  • 第条茞出取匕
    利甚者は、本プログラムおよび凊理結果の利甚に関連しお倖囜為替及び倖囜貿易法これに関連する政省什を含むたたは米囜茞出管理法什で芏定する蚱可が必芁な茞出を行うずきは、利甚者自らが所定の蚱可を取埗するものずしたす。

  • 第条管蜄裁刀所
    本利甚芏玄に関し生じた玛争に぀いおは、東京地方裁刀所をもっお第䞀審の専属的合意管蜄裁刀所ずしたす。

  • 第条準拠法
    本利甚芏玄は日本法に準拠したす。

  • 第条その他の芏定
    本芏玄は、本プログラムの利甚者ず提䟛者ずの間の利甚に関する党おの事項を定めるものであり、本芏玄に定めのない事項に぀いおは、関係法什に埓うものずしたす。

  • 第条蚀語
    本芏玄は日本語を正本ずしたす。本芏玄の英蚳版は、参考のために䜜成されたものであり、䜕らの法的拘束力もないものずしたす。

以䞊

LLM-jp-3 172B beta1 Terms of Use

This Terms of Use (hereinafter referred to as "TOU") sets forth the conditions for the use of the large-scale language model LLM-jp-3 172B beta1 (hereinafter referred to as "the Program") that is made public as a result of the development by the Research and Development Center for Large Language Models at the National Institute of Informatics (hereinafter referred to as "the Provider"). Users of the Program (hereinafter referred to as "Users") shall use the Program upon agreeing to the TOU.

  • Article 1 (License to Use)

    1. Users of the Program must apply for the use of the Program by a method separately specified in addition to the TOU and obtain individual permission from the Provider.
    2. Users may use the Program for commercial or non-commercial purposes in accordance with the TOU. Users are allowed to modify and duplicate the Program, but redistribution of the Program and/or the large-scale language model created by modifying the Program (hereinafter referred to as "Modified Works") is prohibited. Users may provide services using the Program or Modified Works, but such services must not allow third parties to access, download, or obtain the Program or Modified Works directly.
    3. Users who violate the TOU are not allowed to use the Program.
  • Article 2 (Responsibility)

    1. Users agree in advance that the Program is provided “AS IS”, and the Provider makes no warranties, express or implied, regarding the Program, including, but not limited to, its accuracy, completeness, up-to-dateness, and quality, and that the Provider shall not be liable for any damages arising from the use or inability to use the Program.
    2. Users shall compensate for any and all damages suffered by the Provider as a result of the use of the Program and/or the Users' violation of the TOU.
    3. Users shall use the Program at their own responsibility and discretion, and shall handle any disputes arising with third parties in relation to the use of the Program at their own responsibility and expense, and shall indemnify, defend and hold harmless the Provider against all damages and losses without causing any inconvenience to the Provider. Users shall deal with any damages caused by the use of the Program at their own responsibility.
  • Article 3 (Prohibited Actions)
    Users shall not engage in the following actions when using the Program.
    (1) Actions that will or may infringe on the intellectual property rights of the Provider or third parties;
    (2) Actions that will or may infringe on the property, privacy, or portrait rights of the Provider or third parties;
    (3) Actions that discriminate against, defame, insult, or slander the Provider or third parties, promote discrimination against others, or damage the reputation or credibility of others;
    (4) Actions that will or may cause inconvenience or harm to the Provider or third parties;
    (5) Actions that engage in unauthorized legal services and/or provide legal advice from anyone other than a qualified professional;
    (6) Actions that provide financial advice from anyone other than a qualified professional;
    (7) Medical actions, including providing health advice or suggesting treatment methods; and
    (8) Other actions that require permissions or other forms of authorization under laws and regulations.

  • Article 4 (Restrictions)

    1. Users acknowledge that the results of processing using the Program (hereinafter referred to as "Processing Results") may contain falsehoods, biases, content that infringes on the rights of others, or content that does not meet the effectiveness or usefulness expected by Users, and agree to use the Program on the premise that inaccurate or inappropriate Processing Results may cause damage or infringement of rights to Users or third parties and/or ethical concerns. Users shall use the Processing Results after confirming their accuracy, legality, and ethical validity themselves. If the use of the Program, including the Processing Results, by Users cause infringement of the rights of the Users themselves or third parties, the Provider shall not be responsible for any damages, and the Users shall indemnify, defend and hold harmless the Provider against all damages and losses without causing any inconvenience to the Provider.
    2. Users shall use the Processing Results in compliance with the regulations such as laws and regulations in each country and region.
    3. Users shall not use the Processing Results for the actions listed in Article 3 (Prohibited Actions).
  • Article 5 (Ownership of Rights)

    1. Except as expressly provided in the TOU, Users shall not acquire any rights in relation to the Program.
    2. Users will acquire rights newly arising from the creation of Modified Works of the Program, but Users shall use Modified Works in accordance with the TOU.
    3. The Provider shall not assert any rights to the Processing Results.
  • Article 6 (Export Transaction)
    Users shall obtain the necessary permissions themselves when exporting the Program and the Processing Results in relation to their use, where such export requires permissions under the Foreign Exchange and Foreign Trade Act (including related cabinet order and ministerial order) or U.S. export control laws and regulations.

  • Article 7 (Jurisdiction)
    The Tokyo District Court shall have exclusive jurisdiction in the court of the first instance over any disputes arising out of or in connection with the TOU.

  • Article 8 (Governing Law)
    The TOU is governed by and construed in accordance with the laws of Japan.

  • Article 9 (Other Provisions)
    The TOU sets forth the entire agreement as to all matters concerning the use of the Program between the Users and the Provider, and matters not provided for in the TOU shall be governed by the relevant laws and regulations.

  • Article 10 (Governing Language)
    The governing language of the TOU shall be Japanese. The English translation hereof is made for reference purpose only and shall have no effect.

Log in or Sign Up to review the conditions and access this model content.

llm-jp-3-172b-beta1

This repository provides large language models developed by the Research and Development Center for Large Language Models at the National Institute of Informatics.

The development was partially supported by GENIAC.

Checkpoints format: Hugging Face Transformers

Required Libraries and Their Versions

  • torch>=2.3.0
  • transformers>=4.40.1
  • tokenizers>=0.19.1
  • accelerate>=0.29.3
  • flash-attn>=2.5.8

Usage

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("llm-jp/llm-jp-3-172b-beta1")
model = AutoModelForCausalLM.from_pretrained("llm-jp/llm-jp-3-172b-beta1", device_map="auto", torch_dtype=torch.bfloat16)
text = "自然蚀語凊理ずは䜕か"
tokenized_input = tokenizer.encode(text, add_special_tokens=False, return_tensors="pt").to(model.device)
with torch.no_grad():
    output = model.generate(
        tokenized_input,
        max_new_tokens=100,
        do_sample=True,
        top_p=0.95,
        temperature=0.7,
        repetition_penalty=1.05,
    )[0]
print(tokenizer.decode(output))

Model Details

  • Model type: Transformer-based Language Model
  • Total seen tokens: 700B
Params Layers Hidden size Heads Context length
172b 96 12288 96 4096

Tokenizer

The tokenizer of this model is based on huggingface/tokenizers Unigram byte-fallback model. The vocabulary entries were converted from llm-jp-tokenizer v3.0. Please refer to README.md of llm-jp-tokenizer for details on the vocabulary construction procedure (the pure SentencePiece training does not reproduce our vocabulary).

Datasets

Pre-training

The models have been pre-trained using a blend of the following datasets.

Language Dataset Tokens
Japanese Wikipedia 2.6B
Common Crawl 762.8B
WARP/PDF 237.3B
WARP/HTML 2.7B
Kaken 1.8B
English Wikipedia 4.7B
Dolma/CC-head 608.5B
Dolma/C4 181.6B
Dolma/Reddit 83.1B
Dolma/PeS2o 62.9B
Dolma/Gutenberg 5.5B
Dolma/Wiki 3.9B
Code The Stack 114.1B
Chinese Wikipedia 0.8B
Korean Wikipedia 0.3B

Instruction tuning

The models have been fine-tuned on the following datasets.

Language Dataset description
Japanese ichikara-instruction-004-002 A manually constructed Japanese instruction dataset
answer-carefully-001 A manually constructed Japanese instruction dataset focusing on LLMs' safety
databricks-dolly-15k-ja databricks-dolly-15k translated into Japanese using DeepL
oasst1-21k-ja A subset of oasst1 translated into Japanese using DeepL
oasst2-33k-ja A subset of oasst2 translated into Japanese using DeepL
aya-dataset-ja A Japanese subset of aya_dataset
ichikara-instruction-format A small amount of instruction dataset edited from ichikara-instruction, with some constraints on the output format.
English databricks-dolly-15k -
oasst1-21k-en A subset of oasst1
oasst2-33k-en A subset of oasst2
Daring-Anteater -
FLAN We used sampled one.

Risks and Limitations

The models released here are in the early stages of our research and development and have not been tuned to ensure outputs align with human intent and safety considerations.

Send Questions to

llm-jp(at)nii.ac.jp

License

See the LICENSE file.

Model Card Authors

The names are listed in alphabetical order.

Hirokazu Kiyomaru and Takashi Kodama.

Downloads last month
-
Safetensors
Model size
172B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support