Add Wikipedia Persian Dataset - #3629
Merged
Merged
Conversation
pourmand1376
requested review from
Vechtomov,
bitplane,
huu4ontocord,
olliestanley and
sedthh
as code owners
August 2, 2023 09:47
Closed
olliestanley
approved these changes
Aug 2, 2023
olliestanley
added a commit
that referenced
this pull request
Aug 3, 2023
The level of importance of this data is less than Wikipedia. So, I think [this pull request](#3629) should be merged first. I have uploaded the data to [huggingface](https://huggingface.co/datasets/pourmand1376/isna-news) according to Open-assistant's standard. So, it shouldn't need any processing. --------- Co-authored-by: Oliver Stanley <olivergestanley@gmail.com>
|
Hi @pourmand1376, sorry for a slighly off-topic question: could you please share any details on how your friend managed to fine-tune LLaMA on text-only dataset, without instructions? I'm interested in doing the same thing with Belarusian Wikipedia, but so far I've only seen tutorials on how to instruct-tune LLaMA, and Wikipedia articles as such don't contain clearly delimited prompts and responses. Could you please briefly describe the approach? Thanks in advance for any comments. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Currently, the Open-assistant model doesn't support Farsi. This is a text-only dataset to learn Farsi (Persian).
One of my friends fine-tuned LLaMa on this dataset and It could understand Farsi grammar and word usage very well. If the Open-assistant team wants to add support to Farsi, this should be the first step.
I have transformed the dataset into the standard that has been mentioned here and uploaded it to my huggingface account.